Skip to content

Dataset generation API

generate_dataset writes a split dataset to disk. SyntheticGenerator yields samples in memory. Both read the same SyntheticConfig, so a configuration proven on one path transfers to the other.

Everything on this page is importable from the synth_datasets root namespace and needs no torch install. The one exception is SyntheticIterableDataset, documented at the end of this page.

For narrative guides, start with Synthetic data generation; for the annotation objects, backgrounds, degradations, shape families, and writers, see Dataset scenes and outputs API.

Entry points

generate_dataset

generate_dataset

generate_dataset(output_dir: str | Path, num_images: int, fmt: OutputFormat | str = COCO, split_ratios: SplitRatios | None = None, seed: int | None = None, config: SyntheticConfig | None = None, *, overwrite: bool = False, **config_kwargs: Any) -> dict[str, int]

Generate a synthetic dataset on disk and return per-split image counts.

Everything about an image's content — task, class mode, shapes, colors, size — is a :class:SyntheticConfig field, reachable either through config or through config_kwargs. Only fmt and split_ratios, which describe the on-disk layout rather than the pixels, are parameters here.

task and class_mode are config fields with a single owner: task="keypoints" passed here arrives as a config field, so the generator and the writer always read the same task.

Parameters:

Name Type Description Default
output_dir str | Path

Destination directory (created if absent).

required
num_images int

Total number of images to generate across all splits.

required
fmt OutputFormat | str

Output layout: "coco", "yolo" (or an :class:OutputFormat), or any key registered with :func:register_writer.

COCO
split_ratios SplitRatios | None

Train/val/test fractions; defaults to 70/20/10.

None
seed int | None

Seed for reproducible generation; None uses fresh entropy.

None
config SyntheticConfig | None

Full :class:SyntheticConfig. When given, config_kwargs must be empty — a config already says everything they would.

None
overwrite bool

Replace an earlier dataset's files instead of refusing. The writer deletes the paths this run writes (for COCO, each <split>/; for YOLO, images/<split>, labels/<split> and data.yaml) plus every path an earlier run recorded in the .vision-synth.json manifest, so a split an earlier run wrote and this run does not (say, test when this run's ratios leave it empty) is removed too. Nothing else under output_dir is touched, and nothing is deleted through a symlink or outside output_dir. The new dataset is written into a staging directory first and the earlier one is replaced only once it is complete, so a failure while generating leaves the earlier dataset as it was (see :meth:DatasetWriter.write_replacing). Without it, a populated destination is refused, since writing into it would mix the two datasets.

False
**config_kwargs Any

:class:SyntheticConfig fields (task, class_mode, img_size, shapes, colors, …) used to build the config when config is not supplied.

{}

Returns:

Type Description
dict[str, int]

Ordered mapping of split name to the number of images written.

Raises:

Type Description
ValueError

If output_dir holds a staging or backup directory an interrupted overwrite left, if num_images is not a positive integer, if both a config and config_kwargs were supplied, if seed is negative, if no writer is registered for fmt, or if overwrite would delete outside output_dir or through a symlink.

FileExistsError

If overwrite is false and a path the writer would write is already populated. Raised before any sample is generated or any file written.

RuntimeError

If overwrite is true and the swap of the new dataset into place fails. It is rolled back first, by renaming only. When the earlier dataset is fully back, the emptied .vision-synth-backup-* directory is removed; otherwise it is kept and named. A KeyboardInterrupt during the swap is rolled back the same way and re-raised.

Examples:

>>> import tempfile
>>> from synth_datasets import generate_dataset
>>> with tempfile.TemporaryDirectory() as tmp:
...     generate_dataset(tmp, num_images=10, fmt="coco", task="segmentation",
...                      img_size=64, seed=1)
{'train': 7, 'val': 2, 'test': 1}

SyntheticGenerator

SyntheticGenerator

SyntheticGenerator(config: SyntheticConfig)

Draw colored shapes into reproducible :class:Sample objects.

Parameters:

Name Type Description Default
config SyntheticConfig

Generation knobs; see :class:~synth_datasets.core.config.SyntheticConfig.

required

Examples:

>>> import numpy as np
>>> from synth_datasets.core.config import SyntheticConfig
>>> from synth_datasets.core.generator import SyntheticGenerator
>>> gen = SyntheticGenerator(SyntheticConfig(img_size=32))
>>> a = gen.sample(np.random.default_rng(1))
>>> b = gen.sample(np.random.default_rng(1))
>>> bool(np.array_equal(a.image, b.image))
True

Store config and precompute the class vocabulary and keypoint schema this run uses.

The vocabulary is narrowed to config.shapes, the same list every writer declares as its categories/names block, so an annotation's class_id always resolves against the vocabulary written beside it (see :func:~synth_datasets.core.config.class_vocabulary). The schema is resolved once here and stamped onto every landmark-bearing annotation, so a table always travels with the family that produced it.

sample

sample(rng: Generator) -> Sample

Generate one image and its annotations.

Parameters:

Name Type Description Default
rng Generator

Random generator driving object count, shapes, colors, and placement. The background, the clutter, the occluders and the degradations each draw from a child of it rather than from it directly (see :meth:_side_streams), so none of them moves an object and none of them moves another.

required

Returns:

Name Type Description
A Sample

class:Sample with an RGB uint8 image and one annotation per drawn shape.

Raises:

Type Description
RuntimeError

If fewer than min_objects shapes could be placed within the overall attempt budget (num_objects * max_placement_attempts); relax overlap_iou or boundary_tolerance, lower min_objects, or raise max_placement_attempts.

Examples:

>>> import numpy as np
>>> from synth_datasets.core.config import SyntheticConfig
>>> from synth_datasets.core.generator import SyntheticGenerator
>>> gen = SyntheticGenerator(SyntheticConfig(img_size=48, min_objects=1, max_objects=3))
>>> s = gen.sample(np.random.default_rng(7))
>>> 1 <= len(s.annotations) <= 3
True

generate

generate(num_images: int, seed: int | SeedSequence | None = None) -> Iterator[Sample]

Lazily yield num_images samples from a fresh seeded generator.

This is the streaming primitive: only one :class:Sample is materialized at a time, so both in-memory training feeds and writing very large datasets stay within a bounded memory footprint. All samples draw from a single :class:numpy.random.Generator, so a fixed seed yields a reproducible stream.

Parameters:

Name Type Description Default
num_images int

Number of samples to yield.

required
seed int | SeedSequence | None

Integer or independent stream seed for the internal generator; None uses fresh entropy.

None

Yields:

Name Type Description
One Sample

class:Sample per iteration.

Examples:

>>> from synth_datasets.core.config import SyntheticConfig
>>> from synth_datasets.core.generator import SyntheticGenerator
>>> gen = SyntheticGenerator(SyntheticConfig(img_size=32))
>>> samples = list(gen.generate(3, seed=0))
>>> len(samples)
3

Command line

vision-synth generate (the cli extra) calls this function; see command line.

generate

generate(output_dir: str | Path, num_images: int, fmt: str = 'coco', split_ratios: SplitSpec | None = None, seed: int | None = None, overwrite: bool = False, shapes: str | Sequence[str] | None = None, colors: str | Sequence[object] | None = None, **config_fields: Any) -> dict[str, int]

Generate a dataset on disk and return its per-split image counts — the generate command.

Parameters:

Name Type Description Default
output_dir str | Path

Destination directory.

required
num_images int

Total number of images across all splits.

required
fmt str

"coco", "yolo", or any key registered with :func:~synth_datasets.register_writer.

'coco'
split_ratios SplitSpec | None

A {name: fraction} mapping or a (train, val, test) triple; None means 70/20/10.

None
seed int | None

Seed for a reproducible dataset; None uses fresh entropy.

None
overwrite bool

Replace an earlier dataset in output_dir; see :func:~synth_datasets.generate_dataset.

False
shapes str | Sequence[str] | None

Shape names, as a sequence or one comma-separated string such as "duck,camel".

None
colors str | Sequence[object] | None

Colours, as one comma-separated string of names (red, green, blue, any case), one (r, g, b) triple, or a sequence of names and triples. --distractor_colors takes the same spellings, and --distractor_shapes the same as shapes.

None
**config_fields Any

Any other :class:~synth_datasets.SyntheticConfig field, e.g. task="obb".

{}

Returns:

Type Description
dict[str, int]

Ordered mapping of split name to the number of images written.

Examples:

>>> import tempfile
>>> from synth_datasets.cli import generate
>>> with tempfile.TemporaryDirectory() as tmp:
...     generate(tmp, 3, fmt="yolo", img_size=32, seed=0, shapes="duck,camel")
{'train': 2, 'val': 1}

Configuration

SyntheticConfig is the single configuration object. generate_dataset accepts either a full config or its individual fields, but not both in one call.

SyntheticConfig dataclass

SyntheticConfig(img_size: int | tuple[int, int] = 640, min_objects: int = 1, max_objects: int = 10, min_size_ratio: float = 0.1, max_size_ratio: float = 0.3, overlap_iou: float = 0.1, boundary_tolerance: float = 0.05, max_placement_attempts: int = 100, background: ColorLike | Background = (128, 128, 128), degrade: tuple[Degradation, ...] = (), distractors: int = 0, occluders: int = 0, distractor_shapes: tuple[Shape, ...] | None = None, distractor_colors: tuple[Fill, ...] | None = None, rotate: bool = True, asymmetry_jitter: float = 0.0, class_mode: ClassMode = SHAPE, shapes: tuple[Shape, ...] = DEFAULT_SHAPES, colors: tuple[Fill, ...] = DEFAULT_COLORS, task: Task = DETECTION)

Knobs controlling one synthetic image's content.

Parameters:

Name Type Description Default
img_size int | tuple[int, int]

Canvas size in pixels: an int for a square canvas, or a (width, height) pair for a rectangular one. :attr:canvas_size reads either back as (width, height); images come out as (height, width, 3) arrays.

640
min_objects int

Minimum objects drawn per image (inclusive).

1
max_objects int

Maximum objects drawn per image (inclusive).

10
min_size_ratio float

Minimum object size as a fraction of the canvas's shorter side.

0.1
max_size_ratio float

Maximum object size as a fraction of the canvas's shorter side.

0.3
overlap_iou float

Reject a candidate whose IoU with any kept box exceeds this.

0.1
boundary_tolerance float

Max fraction of a box allowed outside the canvas.

0.05
max_placement_attempts int

Retry cap per object before giving up.

100
background ColorLike | Background

What fills the canvas before any object is drawn. Either a plain fill — a :class:Color or its name ("red"), an (r, g, b) triple, or a :class:Fill — or a :class:~synth_datasets.content.backgrounds.Background such as :class:~synth_datasets.content.backgrounds.NoiseBackground or :class:~synth_datasets.content.backgrounds.TextureBackground. A plain fill is normalized to a :class:~synth_datasets.content.backgrounds.SolidBackground at construction, so config.background always reads back as a background object and the default draws a flat grey canvas. Only the three :class:Color names are accepted as strings; other Pillow colour strings such as "white" or "#204080" are rejected — pass an (r, g, b) triple instead. A background renders from its own side stream, never from the placement stream, so switching one on never moves an object at a fixed seed.

(128, 128, 128)
rotate bool

Apply a random rotation to each polygonal shape.

True
asymmetry_jitter float

Max fraction, in [0, 0.5), by which one randomly chosen half of a shape — left or right of its own local vertical axis, before rotation — is narrowed, drawn independently per placed object. 0.0 (the default) disables it and leaves every existing seeded configuration's output unchanged. Every shape this package draws except :attr:~synth_datasets.families.primitives.PrimitiveShape.CIRCLE is mirror-symmetric about that axis in its canonical orientation, so its oriented bounding box would otherwise always show identical left/right margins; a nonzero value breaks that with per-instance variety instead — real oriented objects (vehicles, ships) are rarely that symmetric. circle is always excluded: it never rotates either, so an unrotated skew would bias every instance toward the same absolute image direction rather than varying with a random orientation. Applies to the polygon and, under :attr:Task.KEYPOINTS, the landmark table together, so a shape and its keypoints never drift apart.

0.0
class_mode ClassMode

How classes are derived (see :class:ClassMode).

SHAPE
shapes tuple[Shape, ...]

Shapes the generator may draw, sampled uniformly. Defaults to :data:DEFAULT_SHAPES; pass e.g. (AnimalShape.DUCK, AnimalShape.GIRAFFE) to draw animal silhouettes instead, tuple(AnimalShape) for every animal, tuple(SymbolShape) for every symbol, tuple(LetterShape) for every letter, or (*PrimitiveShape, *AnimalShape, *SymbolShape, *LetterShape) for the full mixed vocabulary. A shape may also be spelled by its name — ("duck", "giraffe"), or a single "duck" — which is how a YAML file or the vision-synth command line can name one; each name is resolved to its member at construction (see :func:~synth_datasets.families.resolve_shape), so config.shapes always reads back as members. distractor_shapes accepts names the same way. Restricting this does renumber classes: the vocabulary a run declares and the ids it stamps both narrow to exactly these shapes, in this order (see :func:class_names), so a symbols-only run numbers its symbols from 0 rather than from their offset into the full :class:Shape enum. Compare runs by class name, not by raw id.

DEFAULT_SHAPES
colors tuple[Fill, ...]

Fills the generator may draw, sampled uniformly. Each may be spelled as a :class:Color member, its name in any case ("red"), a raw (r, g, b) triple, or a :class:Fill; all four are normalized to :class:Fill at construction, so config.colors reads back as Fill objects whichever spelling went in. Defaults to :data:DEFAULT_COLORS (all three named colors); pass e.g. (Color.RED,) to draw only red objects, or ((255, 215, 0),) for a custom yellow. Under :attr:ClassMode.COLOR and :attr:ClassMode.SHAPE_COLOR restricting this does renumber classes, exactly as shapes does: the color axis narrows to exactly these fills, in this order, so colors=(Color.BLUE,) declares one color class, blue, with id 0. Under :attr:ClassMode.SHAPE there is no color axis, so the ids never depend on this.

DEFAULT_COLORS
task Task

Annotation task the generated samples target, as a :class:Task or its string value ("detection", "segmentation", "obb", "keypoints"). This is the only place a run's task is set — :func:~synth_datasets.generate_dataset reads it from here rather than taking its own argument, so the generator and the writer can never disagree about it. Only :attr:Task.KEYPOINTS changes what the generator computes (it adds the landmark table); the other tasks all read the same polygon/box fields, so they differ at write time only.

DETECTION

Raises:

Type Description
ValueError

On non-positive sizes, inverted min/max ranges, an overlap_iou or boundary_tolerance outside [0, 1], max_placement_attempts below 1, an asymmetry_jitter outside [0, 0.5), a shapes tuple that is empty, names an unknown shape (the message lists every valid name), or holds an element that is neither a :class:Shape nor a name, a colors tuple that is empty or holds an element that is no valid fill, a task naming no :class:Task, or a :attr:Task.KEYPOINTS task combined with a shapes tuple that does not belong entirely to one keypoint-bearing family (see :func:keypoint_schema_for) — a :class:~synth_datasets.families.primitives.PrimitiveShape mixed in, or two of :class:~synth_datasets.families.animals.AnimalShape, :class:~synth_datasets.families.symbols.SymbolShape, and :class:~synth_datasets.families.letters.LetterShape mixed together, since only one landmark schema can describe a dataset.

Examples:

>>> from synth_datasets.families.animals import AnimalShape
>>> from synth_datasets.core.config import ClassMode, Color, SyntheticConfig, Task, class_names
>>> SyntheticConfig(img_size=128).img_size
128
>>> SyntheticConfig(shapes=(AnimalShape.DUCK, AnimalShape.CAMEL)).shapes
(<AnimalShape.DUCK: 'duck'>, <AnimalShape.CAMEL: 'camel'>)
>>> SyntheticConfig(task=Task.KEYPOINTS, shapes=(AnimalShape.DUCK,)).task
<Task.KEYPOINTS: 'keypoints'>
>>> SyntheticConfig(colors=(Color.RED,)).colors
(Fill(rgb=(255, 0, 0), name='red'),)
>>> SyntheticConfig(colors=((255, 215, 0),)).colors[0].label
'ffd700'
>>> blue_only = SyntheticConfig(class_mode=ClassMode.COLOR, colors=(Color.BLUE,))
>>> class_names(blue_only.class_mode, blue_only.shapes, blue_only.colors)
['blue']
>>> SyntheticConfig(shapes=("duck", "camel")).shapes
(<AnimalShape.DUCK: 'duck'>, <AnimalShape.CAMEL: 'camel'>)

canvas_size property

canvas_size: tuple[int, int]

Return the canvas as (width, height) in pixels, whichever way :attr:img_size spells it.

Examples:

>>> from synth_datasets import SyntheticConfig
>>> SyntheticConfig(img_size=64).canvas_size, SyntheticConfig(img_size=(96, 48)).canvas_size
((64, 64), (96, 48))

resolved_background property

resolved_background: Background

Return the canvas filler, typed as the :class:Background the field always holds.

:attr:background declares the union its constructor accepts, because a dataclass builds __init__ from the annotation. Everything past __post_init__ holds a background, but a type checker reading the annotation cannot know that, and rejects config.background.render(...) on the three fill members of the union. This property is where that invariant is stated once, so no caller has to assert or cast it.

Examples:

>>> from synth_datasets.core.config import SyntheticConfig
>>> type(SyntheticConfig(img_size=32, background=(10, 20, 30)).resolved_background).__name__
'SolidBackground'

Raises:

Type Description
TypeError

If the field somehow holds a fill, which only a write that bypasses __post_init__ can produce.

resolved_distractor_shapes property

resolved_distractor_shapes: tuple[Shape, ...]

Return the shapes clutter is actually drawn from.

:attr:distractor_shapes when it was given, otherwise the complement of :attr:shapes against :data:~synth_datasets.families.ALL_SHAPES, so clutter never wears the silhouette of a class. Derived on read rather than stored, which is what keeps it correct after a :func:dataclasses.replace that changed shapes.

Examples:

>>> from synth_datasets.core.config import SyntheticConfig
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> config = SyntheticConfig(img_size=32, shapes=(PrimitiveShape.SQUARE,))
>>> PrimitiveShape.SQUARE in config.resolved_distractor_shapes
False

resolved_distractor_colors property

resolved_distractor_colors: tuple[Fill, ...]

Return the fills clutter is actually drawn from.

:attr:distractor_colors when it was given, otherwise :data:DISTRACTOR_PALETTE minus any entry whose RGB triple :attr:colors already claims. Compared on the triple rather than on the whole :class:Fill, since a fill also carries a name and a user may claim a palette colour under one of their own. Derived on read, for the same reason as :attr:resolved_distractor_shapes.

Examples:

>>> from synth_datasets.core.config import DISTRACTOR_PALETTE, Fill, SyntheticConfig
>>> config = SyntheticConfig(img_size=32, colors=(Fill(rgb=DISTRACTOR_PALETTE[0].rgb),))
>>> DISTRACTOR_PALETTE[0] in config.resolved_distractor_colors
False

SplitRatios dataclass

SplitRatios(train: float = 0.7, val: float = 0.2, test: float = 0.1, named: Mapping[str, float] | None = None)

Dataset split fractions; must be non-negative and sum to ~1.

The three standard splits are the constructor's arguments because they are what almost every caller wants. They are not a limit: :meth:custom takes any names at all, so a fourth calibration split or a bare train/test pair needs no change here. The class used to hardcode exactly train/val/test, which meant "train and test only" had to be spelled as val=0.0 and a fourth split was simply impossible.

Parameters:

Name Type Description Default
train float

Training fraction.

0.7
val float

Validation fraction.

0.2
test float

Test fraction.

0.1

Raises:

Type Description
ValueError

If any fraction is negative or they do not sum to 1.

Examples:

>>> from synth_datasets.core.config import SplitRatios
>>> SplitRatios().to_dict()
{'train': 0.7, 'val': 0.2, 'test': 0.1}
>>> SplitRatios(0.8, 0.2, 0.0).to_dict()
{'train': 0.8, 'val': 0.2}
>>> SplitRatios.custom({"train": 0.6, "calib": 0.2, "test": 0.2}).to_dict()
{'train': 0.6, 'calib': 0.2, 'test': 0.2}

custom classmethod

custom(splits: Mapping[str, float]) -> SplitRatios

Build split ratios over arbitrary split names.

Parameters:

Name Type Description Default
splits Mapping[str, float]

name -> fraction, in the order the splits should be written. Fractions must be non-negative and sum to 1, exactly as for the standard three.

required

Returns:

Name Type Description
The SplitRatios

class:SplitRatios carrying those splits.

Raises:

Type Description
ValueError

If splits is empty, a name is not one plain path component (see :func:validate_split_name), or its fractions are negative or do not sum to 1.

Examples:

>>> from synth_datasets.core.config import SplitRatios
>>> SplitRatios.custom({"train": 0.9, "holdout": 0.1}).to_dict()
{'train': 0.9, 'holdout': 0.1}

to_dict

to_dict() -> dict[str, float]

Return the non-zero splits as an ordered name -> fraction mapping.

as_canvas_size

as_canvas_size(img_size: int | tuple[int, int]) -> tuple[int, int]

Return a canvas size as (width, height), validating it.

The one place the two spellings of :attr:SyntheticConfig.img_size are unpacked: a plain int is a square side, a pair is (width, height) — Pillow's order, not numpy's (rows, columns). Every built-in :class:~synth_datasets.content.backgrounds.Background normalizes its img_size argument through here.

Parameters:

Name Type Description Default
img_size int | tuple[int, int]

A positive int, or a pair of positive int values (width, height).

required

Returns:

Type Description
tuple[int, int]

(width, height).

Raises:

Type Description
ValueError

If img_size is not a positive int or a pair of them. bool is refused though it subclasses int: img_size=True is a mistake, not a one-pixel canvas.

Examples:

>>> from synth_datasets.core.config import as_canvas_size
>>> as_canvas_size(64), as_canvas_size((96, 48))
((64, 64), (96, 48))

Tasks, formats, and fills

Task

Bases: str, Enum

Annotation task the dataset targets.

Attributes:

Name Type Description
DETECTION

Axis-aligned bounding boxes only.

SEGMENTATION

Bounding boxes plus filled polygon masks.

OBB

Oriented (rotated) bounding boxes as four corner points.

KEYPOINTS

Bounding boxes plus the named landmarks of whichever keypoint-bearing family the run draws from — animals (16 anatomical points), symbols (7), or letters (15 grid nodes). A run must stay within one such family, since a dataset declares one landmark schema; :func:~synth_datasets.families.keypoint_schema_for is what resolves it. Points a shape does not have (a whale's hind limbs, a grid slot a letter never touches) are absent rather than faked — see each family's module for the NaN-row contract.

Examples:

>>> from synth_datasets.core.config import Task
>>> Task.OBB.value
'obb'
>>> Task("keypoints") is Task.KEYPOINTS
True

OutputFormat

Bases: str, Enum

On-disk dataset layout.

Attributes:

Name Type Description
COCO

Roboflow-style <split>/_annotations.coco.json plus images.

YOLO

images/<split> + labels/<split> + data.yaml.

Fill dataclass

Fill(rgb: tuple[int, int, int], name: str | None = None)

One object fill: the RGB triple to draw with, plus the name it came from when it had one.

The single fill type everything past the config boundary holds. Every accepted spelling — a :class:Color, its name, or a raw (r, g, b) triple — normalizes to this one type, so no consumer takes a union apart again for an RGB or a label.

Parameters:

Name Type Description Default
rgb tuple[int, int, int]

The (r, g, b) triple to draw with; three integers in [0, 255].

required
name str | None

The :class:Color name this fill came from, or None for a raw triple. Only :attr:label reads it.

None

Raises:

Type Description
ValueError

If rgb is not a tuple of three integers in [0, 255].

Examples:

>>> from synth_datasets.core.config import Color, Fill
>>> Fill.parse(Color.BLUE).rgb
(0, 0, 255)
>>> Fill.parse((255, 215, 0)).label
'ffd700'

label property

label: str

Return the class-name label this fill contributes.

A named fill labels itself; a raw triple has no name, so it is labelled by its hex value — (255, 215, 0) becomes "ffd700". That keeps :attr:ClassMode.COLOR and :attr:ClassMode.SHAPE_COLOR well defined for custom fills without inventing color names.

Examples:

>>> from synth_datasets.core.config import Color, Fill
>>> Fill.parse(Color.RED).label, Fill.parse((255, 215, 0)).label
('red', 'ffd700')

parse classmethod

parse(color: ColorLike) -> Fill

Normalize any accepted spelling of a fill into a :class:Fill.

The one boundary where the :data:ColorLike union is unpacked. Every signature that takes a caller-supplied fill runs it through here once and holds the result.

Parameters:

Name Type Description Default
color ColorLike

A :class:Color member, its name in any case ("red", "Blue"), an (r, g, b) tuple, or an existing :class:Fill.

required

Returns:

Type Description
Fill

The normalized fill — color itself when it already is one.

Raises:

Type Description
ValueError

If color is a string naming no :class:Color (the message lists the valid names), or is neither a name, a :class:Color, nor a valid RGB triple.

Examples:

>>> from synth_datasets.core.config import Color, Fill
>>> Fill.parse(Color.BLUE)
Fill(rgb=(0, 0, 255), name='blue')
>>> Fill.parse((255, 215, 0))
Fill(rgb=(255, 215, 0), name=None)
>>> Fill.parse("Blue") == Fill.parse(Color.BLUE)
True

Color

Bases: str, Enum

The named fill vocabulary; each member carries its own 8-bit RGB payload.

A closed set of three, which is what :attr:ClassMode.COLOR numbers its classes from. A run is not limited to them: a fill may equally be a raw (r, g, b) triple, so a yellow object needs no change here. Both spellings normalize to a :class:Fill at the boundary, and that is the single type everything downstream holds; background and object fills accept the same spellings.

The RGB triple is stored on the member itself, through __new__, rather than looked up in a table rebuilt on every access. Members stay ordinary strings either way: Color.RED == "red" holds, Color("red") parses, and JSON serializes a member as its value.

Attributes:

Name Type Description
RED

(255, 0, 0).

GREEN

(0, 128, 0).

BLUE

(0, 0, 255).

Examples:

>>> from synth_datasets.core.config import Color
>>> Color.GREEN.rgb
(0, 128, 0)
>>> Color("red") is Color.RED
True

rgb property

rgb: tuple[int, int, int]

Return the 8-bit RGB fill tuple for this color.

Class vocabularies

class_mode selects which vocabulary the exported class indices follow. COCO categories are 1-based on disk and YOLO classes are 0-based; the helpers below report the 0-based in-memory indices. See Annotation formats for the on-disk conventions.

ClassMode

Bases: str, Enum

How object classes are derived from shape and color.

Attributes:

Name Type Description
SHAPE

One class per shape in the run's own shapes vocabulary — the four primitives by default, up to all 49 across every family.

COLOR

One class per fill in the run's own colors, named by :attr:Fill.label.

SHAPE_COLOR

Cartesian product of the two, named "<color>_<shape>".

ClassEntry dataclass

ClassEntry(index: int, name: str, shape: Shape | None, color: Fill | None)

One class in a run's vocabulary, with the shape and color it was derived from kept intact.

The point of this type is that the writers never have to reconstruct structure from a class name. A ClassMode.SHAPE_COLOR name is built by concatenation ("red_duck"), and the writers used to recover the shape half by splitting on the first underscore — which is correct only for as long as no shape value and no color value contains an underscore. Carrying the pair through instead makes that whole class of bug impossible.

Parameters:

Name Type Description Default
index int

The class id — this entry's position in its vocabulary.

required
name str

The rendered class name, as it appears in a COCO categories block or a YOLO names list.

required
shape Shape | None

The shape this class was derived from, or None under :attr:ClassMode.COLOR, whose classes name no specific shape.

required
color Fill | None

The color this class was derived from, or None under :attr:ClassMode.SHAPE, whose classes name no specific color.

required

Examples:

>>> from synth_datasets.core.config import ClassMode, class_vocabulary
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> entry = class_vocabulary(ClassMode.SHAPE, (PrimitiveShape.SQUARE,)).entries[0]
>>> entry.index, entry.name, entry.color is None
(0, 'square', True)

ClassVocabulary dataclass

ClassVocabulary(class_mode: ClassMode, entries: tuple[ClassEntry, ...])

The ordered classes one run declares, and the (shape, color) -> id map over them.

A dataset's ids are local to its own vocabulary: :class:SyntheticConfig narrows shapes, and both the categories/names block a writer emits and the ids the generator stamps come from this one object. That is what keeps a written dataset internally consistent — at the cost of an id meaning different things in a narrowed run and a full-vocabulary one. Compare two runs by class name, never by raw id.

Parameters:

Name Type Description Default
class_mode ClassMode

The naming rule these entries were built under.

required
entries tuple[ClassEntry, ...]

The classes, in id order.

required

Examples:

>>> from synth_datasets.core.config import ClassMode, Color, class_vocabulary
>>> from synth_datasets.families.animals import AnimalShape
>>> vocab = class_vocabulary(ClassMode.SHAPE, (AnimalShape.DUCK, AnimalShape.CAMEL))
>>> vocab.names
['duck', 'camel']
>>> vocab.id_of(AnimalShape.CAMEL, Color.RED)
1

names property

names: list[str]

Return the class names in id order — the list a writer declares verbatim.

id_of

id_of(shape: Shape, color: ColorLike) -> int

Return the class id for a (shape, color) pair under this vocabulary's naming rule.

Parameters:

Name Type Description Default
shape Shape

The :data:~synth_datasets.families.Shape member drawn.

required
color ColorLike

The fill drawn — a :class:Color member or a raw (r, g, b) triple.

required

Returns:

Type Description
int

The zero-based class id, indexing into :attr:entries.

Raises:

Type Description
KeyError

If the pair names no class in this vocabulary — under a shape-naming mode, that means shape was not among the shapes the vocabulary was built from.

class_vocabulary

class_vocabulary(class_mode: ClassMode, shapes: Iterable[Shape], colors: Iterable[ColorLike] = DEFAULT_COLORS) -> ClassVocabulary

Build the ordered class vocabulary for a class mode over a specific shape vocabulary.

shapes is required rather than defaulted. It used to default to the full 49-shape vocabulary while :class:SyntheticConfig defaulted to the four primitives, so the obvious call for a default config returned a vocabulary twelve times too large — and silently, since the ids still resolved. Requiring the argument removes the mismatch instead of documenting it; pass :data:~synth_datasets.families.ALL_SHAPES for the full vocabulary.

Parameters:

Name Type Description Default
class_mode ClassMode

The selected :class:ClassMode.

required
shapes Iterable[Shape]

The shapes a run draws from, in the order they should be numbered.

required
colors Iterable[ColorLike]

The fills a run draws from, in the order they are numbered. Defaults to the three named :class:Color members. Only the color-naming modes read it.

DEFAULT_COLORS

Returns:

Name Type Description
The ClassVocabulary

class:ClassVocabulary for that combination.

Raises:

Type Description
ValueError

If a factor that defines classes repeats, or two classes would share a name.

Examples:

>>> from synth_datasets.core.config import ClassMode, class_vocabulary
>>> from synth_datasets.families import ALL_SHAPES, DEFAULT_SHAPES
>>> class_vocabulary(ClassMode.SHAPE, DEFAULT_SHAPES).names
['square', 'rectangle', 'triangle', 'circle']
>>> len(class_vocabulary(ClassMode.SHAPE, ALL_SHAPES).entries)
49
>>> class_vocabulary(ClassMode.COLOR, DEFAULT_SHAPES).names
['red', 'green', 'blue']
>>> class_vocabulary(ClassMode.SHAPE_COLOR, DEFAULT_SHAPES).names[:2]
['red_square', 'green_square']

class_names

class_names(class_mode: ClassMode, shapes: Iterable[Shape], colors: Iterable[ColorLike] = DEFAULT_COLORS) -> list[str]

Return the ordered class-name vocabulary for a class mode and shape vocabulary.

Thin convenience over :func:class_vocabulary for callers that only want the names — the list index is the class id, for the same shapes.

Parameters:

Name Type Description Default
class_mode ClassMode

The selected :class:ClassMode.

required
shapes Iterable[Shape]

The shapes a run draws from, in the order they should be numbered.

required
colors Iterable[ColorLike]

The fills a run draws from, in the order they are numbered. Defaults to the three named :class:Color members. Only the color-naming modes read it.

DEFAULT_COLORS

Returns:

Type Description
list[str]

Class names in id order.

Examples:

>>> from synth_datasets.core.config import ClassMode, class_names
>>> from synth_datasets.families import ALL_SHAPES, DEFAULT_SHAPES
>>> class_names(ClassMode.SHAPE, DEFAULT_SHAPES)
['square', 'rectangle', 'triangle', 'circle']
>>> class_names(ClassMode.SHAPE, ALL_SHAPES)[4:8]
['duck', 'elephant', 'giraffe', 'fish']
>>> class_names(ClassMode.SHAPE, ALL_SHAPES)[23:27]
['a', 'b', 'c', 'd']
>>> len(class_names(ClassMode.SHAPE, ALL_SHAPES))
49

class_id

class_id(shape: Shape, color: ColorLike, class_mode: ClassMode, shapes: Iterable[Shape], colors: Iterable[ColorLike] = DEFAULT_COLORS) -> int

Return the class id for a (shape, color) pair under a class mode and shape vocabulary.

Parameters:

Name Type Description Default
shape Shape

The :data:~synth_datasets.families.Shape member.

required
color ColorLike

The :class:Color member.

required
class_mode ClassMode

The selected :class:ClassMode.

required
shapes Iterable[Shape]

The shape vocabulary to index into, in order.

required
colors Iterable[ColorLike]

The fills a run draws from, in the order they are numbered. Defaults to the three named :class:Color members. Only the color-naming modes read it.

DEFAULT_COLORS

Returns:

Type Description
int

Zero-based class id indexing into class_names(class_mode, shapes).

Raises:

Type Description
KeyError

If shape is not in shapes (under a shape-naming mode).

Examples:

>>> from synth_datasets.core.config import ClassMode, Color, class_id
>>> from synth_datasets.families import ALL_SHAPES
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> from synth_datasets.families.symbols import SymbolShape
>>> class_id(PrimitiveShape.TRIANGLE, Color.RED, ClassMode.SHAPE, ALL_SHAPES)
2
>>> class_id(PrimitiveShape.TRIANGLE, Color.RED, ClassMode.COLOR, ALL_SHAPES)
0
>>> class_id(SymbolShape.KITE, Color.RED, ClassMode.SHAPE, ALL_SHAPES)
16
>>> class_id(SymbolShape.KITE, Color.RED, ClassMode.SHAPE, (SymbolShape.KITE,))
0

Coordinate helpers

synth_datasets.families.geometry holds the pixel-space conversions the writers and annotations rely on. Polygons and oriented-box corners are in pixel-centre space; axis-aligned boxes are in pixel-edge space. The two differ by half a pixel each way and are not interchangeable.

to_pixel_centre

to_pixel_centre(points: NDArray[float64]) -> NDArray[float64]

Convert edge-space coordinates to the pixel-centre space the point transforms assume.

Parameters:

Name Type Description Default
points NDArray[float64]

Coordinates in edge space, any shape ending in the coordinate axis.

required

Returns:

Type Description
NDArray[float64]

The same array shifted by - :data:PIXEL_CENTRE_OFFSET.

Examples:

>>> import numpy as np
>>> from synth_datasets.families.geometry import to_pixel_centre
>>> to_pixel_centre(np.array([[3.0, 4.0]]))
array([[2.5, 3.5]])

to_pixel_edge

to_pixel_edge(points: NDArray[float64]) -> NDArray[float64]

Convert pixel-centre coordinates back to the edge space the rasterizer and writers use.

Inverse of :func:to_pixel_centre; the dataset writers apply it so an exported COCO or YOLO file stays in one space with the bbox beside it.

Parameters:

Name Type Description Default
points NDArray[float64]

Coordinates in pixel-centre space, any shape ending in the coordinate axis.

required

Returns:

Type Description
NDArray[float64]

The same array shifted by + :data:PIXEL_CENTRE_OFFSET.

Examples:

>>> import numpy as np
>>> from synth_datasets.families.geometry import to_pixel_edge
>>> to_pixel_edge(np.array([[2.5, 3.5]]))
array([[3., 4.]])

polygon_to_bbox_xyxy

polygon_to_bbox_xyxy(points: NDArray[float64]) -> tuple[float, float, float, float]

Return the axis-aligned bounding box (x_min, y_min, x_max, y_max) of a polygon.

Parameters:

Name Type Description Default
points NDArray[float64]

(num_points, 2) array of coordinates.

required

Returns:

Type Description
tuple[float, float, float, float]

Bounding box tuple in pixels.

Examples:

>>> import numpy as np
>>> from synth_datasets.families.geometry import polygon_to_bbox_xyxy
>>> polygon_to_bbox_xyxy(np.array([[1.0, 2.0], [3.0, 5.0], [0.0, 4.0]]))
(0.0, 2.0, 3.0, 5.0)

polygon_to_obb

polygon_to_obb(points: NDArray[float64], angle: float = 0.0) -> NDArray[float64]

Return the oriented bounding box aligned with the shape's own upright frame.

The box is the polygon's axis-aligned bounding box in the shape's pre-rotation frame, carried rigidly through the placement rotation: the polygon is de-rotated by -angle about its centroid, its min/max extents taken, and the four corners rotated back by angle. Every shape in this package is authored upright (mirror-symmetric about its local vertical axis where it has a symmetry at all — see :mod:~synth_datasets.families.symbols), so the box's sides always run along and across that upright axis, matching how a human would draw the box around the object.

This deliberately is not the minimum-area rectangle. A minimum-area box must lie flush to a convex-hull edge, so for a shape with no horizontal or vertical hull edge in its upright pose (kite, arrow, teardrop) it sits tilted against the shape's own symmetry axis — geometrically tighter, but visually wrong as a pose annotation.

Parameters:

Name Type Description Default
points NDArray[float64]

(num_points, 2) array of polygon coordinates in the image frame.

required
angle float

Rotation in radians that was applied when the polygon was placed (see :func:place_points). Defaults to 0.0, which returns the plain axis-aligned box — a rotated polygon passed without its angle gets a loose upright box, not the tight oriented one, so real callers must pass the placement angle through.

0.0

Returns:

Type Description
NDArray[float64]

(4, 2) array of corner coordinates in order (a rigid rotation of

NDArray[float64]

(min, min) → (max, min) → (max, max) → (min, max) in the de-rotated frame).

Examples:

>>> import numpy as np
>>> from synth_datasets.families.geometry import polygon_to_obb
>>> corners = polygon_to_obb(np.array([[0.0, 0.0], [2.0, 0.0], [2.0, 1.0], [0.0, 1.0]]))
>>> corners.shape
(4, 2)
>>> corners[2] - corners[0]
array([2., 1.])

rotate_polygon

rotate_polygon(points: NDArray[float64], angle: float, center: tuple[float, float] = (0.0, 0.0)) -> NDArray[float64]

Rotate points by angle radians about center.

Parameters:

Name Type Description Default
points NDArray[float64]

(num_points, 2) array of (x, y) coordinates.

required
angle float

Rotation angle in radians (counter-clockwise).

required
center tuple[float, float]

Pivot point (x, y).

(0.0, 0.0)

Returns:

Type Description
NDArray[float64]

Rotated (num_points, 2) array.

Examples:

>>> import numpy as np
>>> from synth_datasets.families.geometry import rotate_polygon
>>> pts = np.array([[1.0, 0.0]])
>>> out = rotate_polygon(pts, np.pi / 2)
>>> bool(np.allclose(out, [[0.0, 1.0]]))
True

bbox_iou

bbox_iou(box_a: tuple[float, float, float, float], box_b: tuple[float, float, float, float]) -> float

Return the intersection-over-union of two axis-aligned xyxy boxes.

Parameters:

Name Type Description Default
box_a tuple[float, float, float, float]

First box (x_min, y_min, x_max, y_max).

required
box_b tuple[float, float, float, float]

Second box (x_min, y_min, x_max, y_max).

required

Returns:

Type Description
float

IoU in [0, 1]; 0.0 when the boxes do not overlap.

Examples:

>>> from synth_datasets.families.geometry import bbox_iou
>>> bbox_iou((0, 0, 2, 2), (1, 1, 3, 3))
0.14285714285714285

PyTorch streaming

SyntheticIterableDataset is the only torch-dependent name in the namespace. It is resolved lazily on first attribute access, so import synth_datasets stays torch-free; using the class requires pip install "vision-synth[torch]". There is no set_epoch() method — construct a new dataset with epoch=... for a fresh stream. See distributed ranks and epochs.

SyntheticIterableDataset

SyntheticIterableDataset(num_images: int, config: SyntheticConfig | None = None, seed: int | None = None, *, rank: int | None = None, world_size: int | None = None, epoch: int = 0, **config_kwargs: Any)

Bases: IterableDataset['Sample']

Stream synthetic :class:Sample objects into a PyTorch training loop.

Parameters:

Name Type Description Default
num_images int

Number of samples yielded by each rank for this epoch.

required
config SyntheticConfig | None

Full :class:SyntheticConfig; when given, config_kwargs are ignored.

None
seed int | None

Base seed for reproducible streams; None uses fresh entropy. Every stream — each rank, epoch, and DataLoader worker — draws from numpy.random.SeedSequence(seed, spawn_key=(rank, epoch, worker_id)), so streams of one seed never collide and neither do streams of adjacent seeds. A single-process stream is therefore not the one :meth:SyntheticGenerator.generate yields for the same seed.

None
rank int | None

Distributed-process rank. None (the default) reads it from the default torch.distributed process group when one is initialized at construction, and means 0 otherwise.

None
world_size int | None

Number of distributed processes. None (the default) reads it from the default torch.distributed process group when one is initialized at construction, and means 1 otherwise.

None
epoch int

Immutable epoch identity for this dataset instance.

0
**config_kwargs Any

Extra :class:SyntheticConfig fields (e.g. img_size, class_mode) used only when config is not supplied.

{}

Examples:

>>> from synth_datasets.export.datasets import SyntheticIterableDataset
>>> ds = SyntheticIterableDataset(num_images=3, img_size=32, seed=1)
>>> samples = list(ds)
>>> len(samples)
3

Store immutable stream identity and build the underlying generator.

rank property

rank: int

Return the immutable distributed rank for this stream.

world_size property

world_size: int

Return the immutable distributed topology size for this stream.

epoch property

epoch: int

Return the immutable epoch identity for this stream.