Dataset scenes and outputs API¶
This page covers what a generated scene is made of and how it reaches disk: the in-memory sample and annotation objects, the four shape vocabularies, the canvas and degradation knobs that set difficulty, and the writers that serialize COCO and YOLO.
For the generation entry points and configuration object, see Dataset generation API. Every knob below is pictured in Customization and extension, and Difficulty bands combines them into three suggested settings.
Samples and annotations¶
SyntheticGenerator.generate yields Sample objects. A sample carries every task representation, so a writer selects the subset its task needs and the generator never needs to know the output format.
Sample
dataclass
¶
Sample(image: NDArray[Any], annotations: list[Annotation], width: int, height: int, scene: SceneRecord = _EMPTY_SCENE)
A rendered image and its annotations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
image
|
NDArray[Any]
|
RGB image, shape |
required |
annotations
|
list[Annotation]
|
Object annotations, one per drawn shape. |
required |
width
|
int
|
Image width in pixels. |
required |
height
|
int
|
Image height in pixels. |
required |
scene
|
SceneRecord
|
What the renderer knows about the scene as a whole — see :class: |
_EMPTY_SCENE
|
Examples:
>>> import numpy as np
>>> from synth_datasets.core.sample import Sample
>>> img = np.zeros((4, 4, 3), dtype=np.uint8)
>>> Sample(img, [], width=4, height=4).width
4
>>> Sample(img, [], width=4, height=4).scene.background_source is None
True
Annotation
dataclass
¶
Annotation(class_id: int, class_name: str, polygon: list[float], bbox_xyxy: tuple[float, float, float, float], angle: float = 0.0, keypoints: tuple[tuple[float, float, int], ...] | None = None, keypoint_schema: KeypointSchema | None = None)
One object instance with all task representations precomputed.
Coordinates are absolute pixel values in the image frame. Polygons and OBB corners are flat
[x1, y1, x2, y2, ...] lists. The oriented box is derived from the polygon on first access
(see :attr:obb_corners) rather than stored, so the tasks that never read it never pay for it.
A landmark table is validated against its own schema on construction (see
:meth:__post_init__), so every consumer — the writers,
:class:~synth_datasets.export.datasets.SyntheticIterableDataset, and any third-party code
reading a :class:Sample — can rely on the width without re-checking it. Validating here rather
than in a writer is what makes that guarantee hold for consumers that never touch a writer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
class_id
|
int
|
Zero-based class index (see :func: |
required |
class_name
|
str
|
Human-readable class label. |
required |
polygon
|
list[float]
|
Filled-shape outline as a flat pixel-coordinate list, in pixel-centre space --
the space :func: |
required |
bbox_xyxy
|
tuple[float, float, float, float]
|
Axis-aligned box |
required |
angle
|
float
|
Rotation in radians the shape was placed with (counter-clockwise, |
0.0
|
keypoints
|
tuple[tuple[float, float, int], ...] | None
|
Landmarks as |
None
|
keypoint_schema
|
KeypointSchema | None
|
The keypoint-bearing family |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> from synth_datasets.families.animals import ANIMAL_KEYPOINT_SCHEMA
>>> from synth_datasets.core.sample import Annotation
>>> ann = Annotation(0, "square", [0.0, 0.0, 2.0, 0.0, 2.0, 2.0, 0.0, 2.0],
... (0.0, 0.0, 2.0, 2.0))
>>> ann.class_name
'square'
>>> ann.keypoints is None
True
>>> table = tuple((1.0, 2.0, 2) for _ in ANIMAL_KEYPOINT_SCHEMA.names)
>>> duck = Annotation(4, "duck", [], (0.0, 0.0, 2.0, 2.0), keypoints=table,
... keypoint_schema=ANIMAL_KEYPOINT_SCHEMA)
>>> duck.keypoints[0]
(1.0, 2.0, 2)
obb_corners
cached
property
¶
Return the upright-frame oriented box as four corners, flat [x1, y1, ..., x4, y4].
The box is the shape's axis-aligned box in its own pre-rotation frame, rotated by
:attr:angle — its sides run along and across the shape's upright (symmetry) axis, not
the minimum-area rectangle's hull-edge direction (see
:func:~synth_datasets.families.geometry.polygon_to_obb).
Derived from :attr:polygon on first access rather than stored, so only runs that read it
pay for the convex hull and rotating-calipers scan; detection, segmentation and keypoint
runs skip that cost entirely.
Returns:
| Type | Description |
|---|---|
list[float]
|
The eight corner coordinates, or an empty list when :attr: |
list[float]
|
three points and has no oriented box to speak of. |
Examples:
SceneRecord
dataclass
¶
What the renderer knows about a scene that its annotations do not say.
The side-car for everything that describes a whole image rather than one object in it. Hanging a
field per feature off :class:Sample would grow its public surface every time the generator
learns something new, and an untyped dict would be unlike every other field in the module;
one typed record means later additions land here and touch no consumer.
eq=False is load-bearing rather than stylistic. A generated __eq__ compares field tuples,
so two populated records would raise ValueError: The truth value of an array with more than one
element is ambiguous instead of returning a bool, and hash() would raise TypeError.
Falling back to identity comparison is well defined for every record and is what a raster
side-car should offer anyway: comparing two masks is the caller's business, not this type's.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
occluder_mask
|
NDArray[bool_] | None
|
|
None
|
background_source
|
str | None
|
Identifier of the image a photographic background was cropped from, or
|
None
|
Examples:
Shape families¶
Four independent vocabularies supply outlines: geometric primitives, traced animal silhouettes, symbols, and stroke letters. Keep one family per keypoint dataset — geometric primitives carry no keypoint schema. See Shape families for the visual reference.
ShapeFamily
dataclass
¶
ShapeFamily(name: str, members: tuple[Shape, ...], base_outline: Callable[[str, float], NDArray[float64]], keypoint_schema: KeypointSchema | None = None, place_keypoints: PlaceKeypoints | None = None)
One shape family's contribution to the drawable vocabulary.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The family's short name, used in error messages and diagnostics ( |
required |
members
|
tuple[Shape, ...]
|
Every member of the family's enum, in declaration order — which is also the order
:func: |
required |
base_outline
|
Callable[[str, float], NDArray[float64]]
|
Returns the origin-centered |
required |
keypoint_schema
|
KeypointSchema | None
|
The family's landmark vocabulary, or |
None
|
place_keypoints
|
PlaceKeypoints | None
|
Places the family's landmark table through the same skew/rotate/translate
pipeline its outline goes through, or |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> from synth_datasets.families import SHAPE_FAMILIES
>>> primitives = SHAPE_FAMILIES[0]
>>> primitives.name, primitives.has_keypoints
('primitives', False)
>>> SHAPE_FAMILIES[1].member_type.__name__
'AnimalShape'
family_of ¶
family_of(shape: Shape) -> ShapeFamily
Return the family a shape member belongs to.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
shape
|
Shape
|
Any :data: |
required |
Returns:
| Type | Description |
|---|---|
ShapeFamily
|
The owning :class: |
Raises:
| Type | Description |
|---|---|
KeyError
|
If |
Examples:
shape_outline ¶
shape_outline(value: str, center: tuple[float, float], size: float, angle: float = 0.0, skew: float = 0.0) -> NDArray[float64]
Build the skewed, rotated, translated outline for any shape value, from any family.
The drawing entry point. It replaced geometry.shape_outline, whose name said "polygon" while
:attr:~synth_datasets.core.sample.Annotation.polygon means the flat coordinate list a
writer emits — two different things one word away from each other.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
A :data: |
required |
center
|
tuple[float, float]
|
Target center |
required |
size
|
float
|
Bounding size in pixels. |
required |
angle
|
float
|
Rotation in radians applied about the shape center. |
0.0
|
skew
|
float
|
Signed fraction narrowing one pre-rotation half — see
:attr: |
0.0
|
Returns:
| Type | Description |
|---|---|
NDArray[float64]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
keypoint_schema_for ¶
keypoint_schema_for(shapes: Iterable[Shape]) -> KeypointSchema | None
Return the schema shared by every shape in shapes, when there is exactly one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
shapes
|
Iterable[Shape]
|
The shapes a run draws from — typically
:attr: |
required |
Returns:
| Name | Type | Description |
|---|---|---|
The |
KeypointSchema | None
|
class: |
KeypointSchema | None
|
|
|
KeypointSchema | None
|
func: |
Examples:
>>> from synth_datasets.families.animals import AnimalShape
>>> from synth_datasets.families import keypoint_schema_for
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> keypoint_schema_for((AnimalShape.DUCK, AnimalShape.CAMEL)).kpt_shape
16
>>> keypoint_schema_for((PrimitiveShape.SQUARE,)) is None
True
Per-family shape enums¶
PrimitiveShape ¶
Bases: ShapeEnum
Analytically computed shape vocabulary (definition order is the class order).
Computed from size rather than looked up in a table. RECTANGLE is deliberately
non-square and every shape but CIRCLE takes a per-shape rotation, so oriented bounding
boxes carry real orientation variety; CIRCLE is rotation-invariant, so its OBB collapses to
the axis-aligned box. None of them carries a landmark table: a square is 4-fold symmetric and a
circle rotation-invariant, so a fixed landmark on them has no identity a model could learn.
Attributes:
| Name | Type | Description |
|---|---|---|
SQUARE |
Axis-aligned equal-sided quadrilateral. |
|
RECTANGLE |
Non-square quadrilateral. |
|
TRIANGLE |
Equilateral triangle, apex up — mirror-symmetric about its vertical axis like every other upright shape in the package. Its 3-fold rotational symmetry means the silhouette alone cannot distinguish rotations 120 degrees apart; the exported oriented box still records the drawn angle exactly. |
|
CIRCLE |
Polygon-approximated circle. |
Examples:
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> [shape.value for shape in PrimitiveShape]
['square', 'rectangle', 'triangle', 'circle']
AnimalShape ¶
Bases: ShapeEnum
Animal silhouette vocabulary (definition order is the animal class order).
Twelve fixed side-profile silhouettes traced from public-domain reference art. Each is
asymmetric and belongs to a distinct silhouette archetype, so the classes stay separable at a
glance and every outline point keeps an unambiguous identity under rotation — the property a
landmark needs and a square or circle cannot offer. Every member has a <value>.svg document
in the packaged assets/animals directory and therefore a landmark table, which is what makes
:attr:~synth_datasets.core.config.Task.KEYPOINTS well-defined for exactly this enum.
Attributes:
| Name | Type | Description |
|---|---|---|
DUCK |
Compact duck silhouette with an S-curved neck and a beak. |
|
ELEPHANT |
Bulky elephant silhouette with a trunk, a large ear, and thick legs. |
|
GIRAFFE |
Tall, thin giraffe silhouette with a very long neck and thin legs. |
|
FISH |
Streamlined fish silhouette with a forked tail fin. |
|
RABBIT |
Compact rabbit silhouette with long upright ears. |
|
CAMEL |
Humped camel silhouette on four long legs. |
|
EAGLE |
Perched eagle silhouette with a hooked beak and a long tail. |
|
PENGUIN |
Upright penguin silhouette with flippers and webbed feet. |
|
WHALE |
Streamlined whale silhouette with a pectoral flipper and a tail fluke. |
|
KANGAROO |
Hopping kangaroo silhouette with a heavy tail and one large hind foot. |
|
FLAMINGO |
Long-legged flamingo silhouette with an S-curved neck. |
|
CROCODILE |
Low, elongated crocodile silhouette with a long snout and sprawled legs. |
Examples:
>>> from synth_datasets.families.animals import AnimalShape
>>> len(AnimalShape)
12
>>> AnimalShape("duck")
<AnimalShape.DUCK: 'duck'>
SymbolShape ¶
Bases: ShapeEnum
Analytic symbol vocabulary (definition order is the symbol class order).
Seven straight-edge 2D symbols, each mirror-symmetric about its own vertical axis and belonging
to a distinct silhouette archetype, so — like :class:~synth_datasets.families.animals.AnimalShape
— every outline point keeps an unambiguous identity under rotation. There is no plain
TRIANGLE/ISOSCELES_TRIANGLE member here: it would collide in name with
:attr:~synth_datasets.families.primitives.PrimitiveShape.TRIANGLE for a shape this family
does not need to keep.
Attributes:
| Name | Type | Description |
|---|---|---|
KITE |
Diamond quadrilateral with unequal top/bottom diagonal lengths. |
|
TRAPEZOID |
Isosceles trapezoid, short parallel side up. |
|
HOUSE |
Square body with a triangular roof (five-sided, convex). |
|
ARROW |
Up-pointing arrow with two barbs (concave). |
|
CROSS |
Latin cross with an elongated lower arm (concave). |
|
TEARDROP |
Rounded top tapering to a bottom point. |
|
ANCHOR |
Ring, stock, shaft and two flukes (concave). |
Examples:
>>> from synth_datasets.families.symbols import SymbolShape
>>> len(SymbolShape)
7
>>> SymbolShape("kite")
<SymbolShape.KITE: 'kite'>
LetterShape ¶
Bases: ShapeEnum
Capital-letter outline vocabulary (definition order is the letter class order).
Twenty-six capitals, each a single simple polygon derived from a small straight-stroke graph on
the shared node grid described in the module docstring. Lower-case only in A-Z.
Examples:
>>> from synth_datasets.families.letters import LetterShape
>>> len(LetterShape)
26
>>> LetterShape("z")
<LetterShape.Z: 'z'>
Keypoint schemas¶
KeypointSchema
dataclass
¶
KeypointSchema(names: tuple[str, ...], skeleton: tuple[tuple[int, int], ...], flip_idx: tuple[int, ...], shape_values: tuple[str, ...], skeleton_by_value: Mapping[str, tuple[tuple[int, int], ...]] | None = None)
The fixed landmark vocabulary one keypoint-bearing shape family draws its tables from.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
names
|
tuple[str, ...]
|
Landmark names, in the order every table, annotation, and label row uses. |
required |
skeleton
|
tuple[tuple[int, int], ...]
|
Visualization-only edges as index pairs into |
required |
flip_idx
|
tuple[int, ...]
|
The landmark each index becomes under a horizontal flip — Ultralytics'
|
required |
shape_values
|
tuple[str, ...]
|
Every :class: |
required |
skeleton_by_value
|
Mapping[str, tuple[tuple[int, int], ...]] | None
|
Per- |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> from synth_datasets.core.keypoints import KeypointSchema
>>> schema = KeypointSchema(names=("a", "b"), skeleton=((0, 1),), flip_idx=(1, 0), shape_values=("x",))
>>> schema.kpt_shape
2
>>> schema.skeleton_for("x")
((0, 1),)
kpt_shape
property
¶
Return the landmark count (Ultralytics' kpt_shape first dimension).
skeleton_for ¶
Return the skeleton edges to draw for one shape_values member.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
shape_value
|
str
|
A value from :attr: |
required |
Returns:
| Type | Description |
|---|---|
tuple[int, int]
|
|
...
|
entry for it; :attr: |
tuple[tuple[int, int], ...]
|
pre- |
Backgrounds¶
background picks the canvas the objects land on. Each background draws from a side stream of its own rather than from the placement stream, so switching one on at a fixed seed cannot move an object.
Background ¶
Bases: ABC
A canvas filler whose subclass is the mode, so there is no mode field.
Implement :meth:render and, when the mode draws nothing, override :attr:consumes_randomness — the generator
reads it to decide whether to take a side stream at all, and a background that claims to draw nothing is handed
None in place of one.
consumes_randomness
property
¶
Return whether :meth:render draws from the generator it is handed.
Defaults to True, which is always safe: a side stream that is taken and not used costs one cheap child and
moves nothing, while a background that draws from a stream it said it would not need would receive None and
fail loudly rather than quietly.
render
abstractmethod
¶
Return the canvas this background fills, as (height, width, 3) uint8.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rng
|
Generator | None
|
The side stream to draw from, or |
required |
img_size
|
CanvasSize
|
The canvas side as an |
required |
Returns:
| Type | Description |
|---|---|
NDArray[uint8]
|
A writable, C-contiguous RGB canvas. |
render_with_source ¶
render_with_source(rng: Generator | None, img_size: CanvasSize) -> tuple[NDArray[uint8], str | None]
Return the canvas together with whatever names where its pixels came from.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rng
|
Generator | None
|
The side stream, exactly as :meth: |
required |
img_size
|
CanvasSize
|
The canvas size, exactly as :meth: |
required |
Returns:
| Type | Description |
|---|---|
NDArray[uint8]
|
The canvas and a provenance string, or |
str | None
|
nothing outside the process to name. |
This is what the generator calls, and what ends up in
:attr:~synth_datasets.core.sample.SceneRecord.background_source. It is not abstract: a
procedural mode has no source, so the default delegates to :meth:render and reports
None, which means a third-party background implementing only :meth:render keeps
working. Only a mode reading files outside the package overrides it.
SolidBackground
dataclass
¶
Bases: Background
One flat colour across the whole canvas — the behaviour that predates this module.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
color
|
ColorLike
|
The fill, as a :class: |
DEFAULT_BASE
|
Examples:
>>> from synth_datasets.content.backgrounds import SolidBackground
>>> SolidBackground((20, 30, 40)).render(None, 2)[0, 0].tolist()
[20, 30, 40]
consumes_randomness
property
¶
Return False: a flat fill has nothing to draw.
render ¶
Return a canvas of one repeated colour, ignoring rng entirely.
Built with :func:numpy.full rather than a broadcast view made contiguous: at img_size=1 every axis is
length one, so the broadcast result is already flagged C-contiguous and :func:numpy.ascontiguousarray hands
it straight back — read-only, which breaks the writability half of the base class's contract at exactly one
canvas size.
GradientBackground
dataclass
¶
GradientBackground(stops: tuple[ColorLike, ColorLike] = ((64, 64, 64), (192, 192, 192)), direction: float | None = None, radial: bool = False)
Bases: Background
A linear or radial ramp between two stops.
Breaks the "background is one value" assumption a colour-mode classifier can otherwise lean on: the same object fill now sits on a different local intensity depending on where it landed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
stops
|
tuple[ColorLike, ColorLike]
|
The two ends of the ramp. For a linear ramp the first is at the low end of the direction axis; for a radial one it is at the canvas centre and the second at the corners. |
((64, 64, 64), (192, 192, 192))
|
direction
|
float | None
|
Ramp angle in radians for a linear ramp, or |
None
|
radial
|
bool
|
Ramp outward from the centre rather than across the canvas. |
False
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> from synth_datasets.content.backgrounds import GradientBackground
>>> ramp = GradientBackground(direction=0.0).render(None, 4)
>>> bool(ramp[0, 0, 0] < ramp[0, -1, 0])
True
consumes_randomness
property
¶
Return whether the ramp angle still has to be sampled.
Only a linear ramp with no fixed direction draws. A radial ramp is centred and has no angle at all, so it
draws nothing whatever direction says — which is why the draw count is keyed on this property rather than on
direction alone.
render ¶
Return the ramp, interpolating between the stops in float and rounding once.
NoiseBackground
dataclass
¶
Bases: Background
Per-pixel Gaussian noise around a base colour.
The first knob that makes a small object genuinely hard: it removes the trivial edge detector a flat canvas hands a model for free.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base
|
ColorLike
|
The colour the noise is centred on. |
DEFAULT_BASE
|
sigma
|
float
|
Per-channel standard deviation in 8-bit units. Keep it at or below
:data: |
16.0
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.backgrounds import NoiseBackground
>>> NoiseBackground(sigma=4.0).render(np.random.default_rng(0), 4).shape
(4, 4, 3)
render ¶
Return the base colour plus one Gaussian field, added rather than multiplied.
ImpulseNoiseBackground
dataclass
¶
ImpulseNoiseBackground(base: ColorLike = DEFAULT_BASE, amount: float = 0.05, salt_ratio: float = 0.5)
Bases: Background
Salt-and-pepper pixels scattered over a base colour.
The same axis as :class:NoiseBackground but heavy-tailed, and the reason it is a separate
difficulty step: impulse pixels survive a blur that erases a Gaussian field.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base
|
ColorLike
|
The colour the surviving pixels keep. |
DEFAULT_BASE
|
amount
|
float
|
Fraction of pixels replaced, in |
0.05
|
salt_ratio
|
float
|
Of those, the fraction set white rather than black, in |
0.5
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.backgrounds import ImpulseNoiseBackground
>>> canvas = ImpulseNoiseBackground(amount=0.5).render(np.random.default_rng(0), 16)
>>> bool((canvas == 255).any() and (canvas == 0).any())
True
render ¶
Return the base colour with a fraction of pixels replaced outright by black or white.
Replacement, not addition: adding a fixed salt value to an arbitrary base would not produce endpoint pixels, which is the whole point of impulse noise and the reason it survives a blur.
TextureBackground
dataclass
¶
TextureBackground(base: ColorLike = DEFAULT_BASE, amplitude: float = 48.0, frequency: float = 8.0, octaves: int = 3, quantize: int | None = None)
Bases: Background
Value noise at a chosen spatial frequency, optionally posterized.
The first background that puts structure at object scale, so a false positive becomes possible rather than merely a matter of contrast.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base
|
ColorLike
|
The colour the field deviates around. |
DEFAULT_BASE
|
amplitude
|
float
|
Maximum deviation from |
48.0
|
frequency
|
float
|
Lattice cells across the image at the first octave. |
8.0
|
octaves
|
int
|
Number of octaves summed, each at twice the frequency and half the weight. |
3
|
quantize
|
int | None
|
Number of levels the normalized |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.backgrounds import TextureBackground
>>> TextureBackground(octaves=2).render(np.random.default_rng(0), 32).shape
(32, 32, 3)
render ¶
Return the base colour plus the summed octaves, added rather than multiplied.
Additive on purpose: a multiplicative field would make the effective contrast depend on
base, so the same amplitude would mean different things on a dark and a light canvas.
ImageBackground
dataclass
¶
Bases: Background
Random crops of the caller's own photographs.
The only mode that reads the outside world, and the only one whose parameter has no default: there is nothing to ship a default directory from, and a silently empty one is the failure worth refusing outright rather than rendering as black.
What it buys is real texture statistics — the spatial correlations, gradients and clutter of
photographs — without any labelling cost, since the labels still come from the shapes drawn on
top. Which file and which crop were used is reported through
:attr:~synth_datasets.core.sample.SceneRecord.background_source, so a sample can be traced
back to what it stood on.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
image_dir
|
Path
|
Directory of images to crop from. Required; scanned and validated at construction. |
required |
grayscale
|
bool
|
Drop the colour of the crop, leaving its structure. Useful when the run's classes are colour-named and a photographic canvas would otherwise compete with them. |
False
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from pathlib import Path
>>> from PIL import Image
>>> import tempfile
>>> with tempfile.TemporaryDirectory() as folder:
... Image.fromarray(np.full((40, 40, 3), 90, np.uint8)).save(Path(folder) / "a.png")
... canvas, source = ImageBackground(Path(folder)).render_with_source(np.random.default_rng(0), 16)
>>> canvas.shape, source
((16, 16, 3), 'a.png')
render ¶
Return one random crop, discarding which file it came from.
render_with_source ¶
render_with_source(rng: Generator | None, img_size: CanvasSize) -> tuple[NDArray[uint8], str | None]
Return one random crop and the name of the file it was taken from.
Three draws, always in this order and always all three: the file index, then the crop's left and top offsets. The offsets are drawn even when only one position is possible, so the draw count does not depend on how large the chosen file happens to be.
Baked degradations¶
degrade bakes camera effects into the exported pixels. Generation applies them once; it does not resample them at each training step.
Degradation ¶
Bases: ABC
One pointwise effect applied to a finished image, in the order the tuple lists it.
Implement :meth:apply and, when the effect has no randomness of its own, override :attr:consumes_randomness —
every parameter here is a fixed scalar by design, so most effects draw nothing and only the ones sampling a field
per pixel do.
consumes_randomness
property
¶
Return whether :meth:apply draws from the generator it is handed.
Defaults to False, the opposite of a background's default: a degradation is configured by explicit scalars,
so drawing is the exception rather than the rule.
apply
abstractmethod
¶
Return the degraded image.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
image
|
NDArray[uint8]
|
|
required |
rng
|
Generator | None
|
The side stream, or |
required |
Returns:
| Type | Description |
|---|---|
NDArray[uint8]
|
A |
NDArray[uint8]
|
build theirs with :func: |
NDArray[uint8]
|
an |
NDArray[uint8]
|
chain ending on one would hand back a sample whose writability depended on which effect |
NDArray[uint8]
|
happened to run last. |
GaussianBlur
dataclass
¶
Bases: Degradation
A Gaussian blur of a fixed radius.
The step that costs corner keypoints and, on a square-ish shape, the oriented box's angle: a blurred corner is no longer a corner, and the pose has to be inferred from the silhouette.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
radius
|
float
|
Blur radius in pixels, as Pillow's :class: |
1.0
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.degradations import GaussianBlur
>>> edge = np.zeros((8, 8, 3), np.uint8)
>>> edge[:, 4:] = 255
>>> bool(GaussianBlur(radius=1.5).apply(edge, None)[0, 3, 0] > 0)
True
apply ¶
Blur through Pillow, which owns the kernel this package does not reimplement.
GaussianNoise
dataclass
¶
Bases: Degradation
Additive per-pixel Gaussian noise over the finished image.
Costs a model edge localisation and small-object recall: the noise floor sits on the object and the canvas alike, so a small shape's boundary stops being the strongest local gradient.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sigma
|
float
|
Per-channel standard deviation in 8-bit units. |
8.0
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
JPEG
dataclass
¶
Bases: Degradation
A JPEG encode/decode round-trip at a fixed quality.
Block artefacts around a thin letter stroke are the point, and they are realistic rather than synthetic: anything that came out of a camera pipeline carries them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
quality
|
int
|
Encoder quality from 1 to 95. Pillow treats anything above 95 as wasted file size rather than added fidelity, so that is the ceiling. |
75
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.degradations import JPEG
>>> JPEG(quality=30).apply(np.full((16, 16, 3), 128, np.uint8), None).shape
(16, 16, 3)
apply ¶
Encode to an in-memory JPEG and decode it back, keeping whatever the codec did.
Contrast
dataclass
¶
Bases: Degradation
A contrast scaling around the image's own mean luminance.
Under :attr:~synth_datasets.core.config.ClassMode.COLOR this is the knob that makes classes
converge toward each other, since every fill moves toward the same grey.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
factor
|
float
|
|
0.7
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.degradations import Contrast
>>> flat = np.stack([np.full((4, 4), v, np.uint8) for v in (0, 255, 128)], axis=2)
>>> int(Contrast(factor=0.0).apply(flat, None).std())
0
apply ¶
Scale every channel toward the image's mean grey, which is what Pillow's enhancer does.
ColorCast
dataclass
¶
Bases: Degradation
A per-channel gain, the way a white balance error looks.
What it costs is red-versus-green separation: two fills that were far apart in hue sit closer together once one channel has been lifted and another pulled down.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gain
|
tuple[float, float, float]
|
One multiplier per channel, in |
(1.1, 1.0, 0.9)
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.degradations import ColorCast
>>> grey = np.full((2, 2, 3), 100, np.uint8)
>>> ColorCast(gain=(1.2, 1.0, 0.8)).apply(grey, None)[0, 0].tolist()
[120, 100, 80]
apply ¶
Multiply each channel by its own gain and clip once.
Vignette
dataclass
¶
Bases: Degradation
A radial darkening toward the corners.
Interacts with boundary_tolerance on purpose: an object allowed to sit half off the frame now
also sits in the darkest part of it, which is the combination a border-region detector fails on.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
strength
|
float
|
Fraction of brightness removed at the corners; |
0.3
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.degradations import Vignette
>>> flat = np.full((9, 9, 3), 200, np.uint8)
>>> out = Vignette(strength=0.5).apply(flat, None)
>>> bool(out[0, 0, 0] < out[4, 4, 0])
True
apply ¶
Scale each pixel by a quadratic falloff from the centre to the corners.
Quantize
dataclass
¶
Bases: Degradation
A reduction to a fixed number of evenly spaced levels per channel.
A flat fill becomes banded, which is the cheap stand-in for a low-bit-depth sensor and the one degradation that makes a gradient background visibly wrong rather than merely dimmer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
levels
|
int
|
Number of levels per channel, from 2 to 256. |
16
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from synth_datasets.content.degradations import Quantize
>>> ramp = np.arange(256, dtype=np.uint8).reshape(16, 16)[..., None].repeat(3, axis=2)
>>> len(np.unique(Quantize(levels=4).apply(ramp, None)))
4
apply ¶
Snap each channel to the nearest of levels evenly spaced values.
Writers¶
A writer turns samples into files. register_writer adds a format under a new name; get_writer resolves one. See Customization and extension for a worked custom writer.
DatasetWriter ¶
DatasetWriter(task: Task, vocabulary: ClassVocabulary, keypoint_schema: KeypointSchema | None = None)
Bases: ABC
Base class for dataset serializers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
task
|
Task
|
The annotation task determining which fields are emitted. |
required |
vocabulary
|
ClassVocabulary
|
The classes to declare, in id order — each entry keeps the shape and color it was derived from, which is how a writer tells which categories its keypoint schema covers without parsing their names. |
required |
keypoint_schema
|
KeypointSchema | None
|
The keypoint family a :attr: |
None
|
Store the task and vocabulary, rejecting a keypoints task with no schema to write.
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
schema
property
¶
schema: KeypointSchema
Return the keypoint schema, which the constructor guarantees for a keypoints task.
Every landmark-writing path reaches the schema through here rather than through the optional attribute, so the "a keypoints writer always has one" invariant is stated once and checked, instead of being asserted implicitly at four call sites.
Raises:
| Type | Description |
|---|---|
ValueError
|
If no schema was supplied — only reachable by mutating the attribute after construction, since the constructor rejects a keypoints task without one. |
owned_paths ¶
Return the paths under output_dir this writer creates for split_names — a subclass hook.
:meth:prepare_output refuses to write while any of them is populated, and :meth:write_replacing swaps
them out, together with whatever an earlier run recorded in the :data:MANIFEST_NAME manifest and nothing
else under output_dir. The base declares none, so a third-party writer that does not override this is
neither blocked nor replaced: it keeps whatever behavior it had.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
split_names
|
Iterable[str]
|
The splits about to be written, each already checked to be one plain path component. |
required |
output_dir
|
Path
|
The destination root. |
required |
Returns:
| Type | Description |
|---|---|
tuple[Path, ...]
|
Every file or directory the writer produces for those splits. |
record_output ¶
Record in output_dir's :data:MANIFEST_NAME manifest the paths this writer now owns there.
The built-in writers call this once a write completes. The manifest keeps an earlier run's entries that still
exist (a split this run left alone is still that tool's own) and adds this run's :meth:owned_paths, so a
later overwrite can remove a split this run did not write without guessing from file names.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
split_names
|
Iterable[str]
|
The splits just written. |
required |
output_dir
|
str | Path
|
The destination root. |
required |
prepare_output ¶
Make sure writing split_names under output_dir cannot mix two datasets; it never deletes anything.
Refuses while any of this run's :meth:owned_paths already holds files, since writing img_000000.. over
a larger earlier run would keep its higher-numbered files beside the new ones. The built-in writers call this
before consuming a single sample; to replace an earlier dataset, use :meth:write_replacing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
split_names
|
Iterable[str]
|
The splits about to be written. |
required |
output_dir
|
str | Path
|
The destination root. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If a split name is not one plain path component (see
:func: |
FileExistsError
|
If an owned path is a file or a non-empty directory. |
write_replacing ¶
write_replacing(splits: dict[str, Iterable[Sample]], output_dir: str | Path) -> None
Write splits over an earlier dataset under output_dir, removing it only once the new one is done.
What overwrite=True runs. The new dataset is written by :meth:write into a fresh
.vision-synth-staging-* directory inside output_dir. Only after that succeeds are the paths to replace —
this run's :meth:owned_paths, every path the earlier :data:MANIFEST_NAME manifest lists, and the manifest
itself — checked (inside output_dir, never through a symlink) and swapped: each is renamed into a fresh
.vision-synth-backup-* directory, then each staged entry is renamed into place, all on one filesystem.
Only after every staged entry is confirmed at its target does the overwrite commit, by renaming the backup and
staging directories to .vision-synth-discard-*; they are removed after that. Both carry a
:data:OWNER_NAME marker, and the backup's records the relative paths the swap replaces and adds.
Any exception during the swap, KeyboardInterrupt included, rolls it back by renaming only, reading the
state from disk. If a re-listing shows the earlier dataset fully back, the emptied backup directory is removed
with os.rmdir alone; otherwise it is kept, whatever it holds, and the error names it. A staging or backup
directory left by an interrupted run makes this refuse (see :func:_refuse_leftovers); it is never removed
automatically.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splits
|
dict[str, Iterable[Sample]]
|
Mapping of split name to its samples; consumed as :meth: |
required |
output_dir
|
str | Path
|
The destination root (created if absent). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If a split name is not one plain path component, the manifest cannot be read, an interrupted
run left a staging or backup directory, a path to replace has a reserved |
RuntimeError
|
If the swap fails (after its rollback; the message names the kept backup and says whether the earlier dataset is fully back), or the new dataset is in place but a leftover cannot be removed. |
write
abstractmethod
¶
write(splits: dict[str, Iterable[Sample]], output_dir: str | Path) -> None
Write all splits under output_dir.
Consume each split exactly once, in the order given. The splits
:func:~synth_datasets.generate_dataset passes are lazy views over a single
shared sample stream, so iterating them out of order, twice, or partially does not merely
repeat work — it silently redistributes samples between splits or empties them. An
implementation that needs a split more than once must materialize it itself, accepting the
memory that costs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splits
|
dict[str, Iterable[Sample]]
|
Mapping of split name to its samples, in the order they must be consumed. |
required |
output_dir
|
str | Path
|
Destination root directory (created if absent). |
required |
CocoWriter ¶
CocoWriter(task: Task, vocabulary: ClassVocabulary, keypoint_schema: KeypointSchema | None = None)
Bases: DatasetWriter
Write a COCO-format dataset, one JSON per split.
Examples:
>>> from synth_datasets.core.config import ClassMode, Task, class_vocabulary
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> from synth_datasets.export.writers import CocoWriter
>>> vocab = class_vocabulary(ClassMode.SHAPE, (PrimitiveShape.SQUARE,))
>>> CocoWriter(Task.DETECTION, vocab).task.value
'detection'
owned_paths ¶
Return each split's <output_dir>/<split>/ directory — images and its JSON both live there.
write ¶
write(splits: dict[str, Iterable[Sample]], output_dir: str | Path) -> None
Stream each split to <output_dir>/<split>/ in a single pass over its samples.
Image pixels are written as they are produced and never held in memory. The COCO schema,
however, emits one JSON document per split, so lightweight per-image and per-annotation
metadata records (no pixels) accumulate for the duration of the split and are serialized
once the split is exhausted: memory is O(n) in the split's image and annotation counts, not
constant. For a constant-memory path use the YOLO writer (one label file per image) or the
in-memory :class:~synth_datasets.export.datasets.SyntheticIterableDataset.
Raises:
| Type | Description |
|---|---|
ValueError
|
If a split name is not one plain path component. |
FileExistsError
|
If a split directory already holds files; see :meth: |
YoloWriter ¶
YoloWriter(task: Task, vocabulary: ClassVocabulary, keypoint_schema: KeypointSchema | None = None)
Bases: DatasetWriter
Write a YOLO-format dataset with normalized labels and a data.yaml.
Examples:
>>> from synth_datasets.core.config import ClassMode, Task, class_vocabulary
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> from synth_datasets.export.writers import YoloWriter
>>> vocab = class_vocabulary(ClassMode.SHAPE, (PrimitiveShape.SQUARE,))
>>> YoloWriter(Task.OBB, vocab).task.value
'obb'
owned_paths ¶
Return each split's images/<split> and labels/<split> directories, plus the shared data.yaml.
write ¶
write(splits: dict[str, Iterable[Sample]], output_dir: str | Path) -> None
Write images, labels, and data.yaml under output_dir.
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
FileExistsError
|
If a split's image or label directory, or |
get_writer ¶
get_writer(fmt: OutputFormat | str, task: Task, vocabulary: ClassVocabulary, keypoint_schema: KeypointSchema | None = None) -> DatasetWriter
Return the writer registered for an output format.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fmt
|
OutputFormat | str
|
Target format — an :class: |
required |
task
|
Task
|
Annotation task to emit. |
required |
vocabulary
|
ClassVocabulary
|
The classes to declare, in id order; see :class: |
required |
keypoint_schema
|
KeypointSchema | None
|
The keypoint family a :attr: |
None
|
Returns:
| Type | Description |
|---|---|
DatasetWriter
|
A concrete :class: |
Raises:
| Type | Description |
|---|---|
ValueError
|
If no writer is registered for |
Examples:
>>> from synth_datasets.core.config import ClassMode, OutputFormat, Task, class_vocabulary
>>> from synth_datasets.families.primitives import PrimitiveShape
>>> from synth_datasets.export.writers import get_writer
>>> vocab = class_vocabulary(ClassMode.SHAPE, (PrimitiveShape.SQUARE,))
>>> type(get_writer(OutputFormat.YOLO, Task.DETECTION, vocab)).__name__
'YoloWriter'
register_writer ¶
register_writer(fmt: OutputFormat | str, writer: type[DatasetWriter]) -> None
Register the writer class serving one output format.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fmt
|
OutputFormat | str
|
The format key. An :class: |
required |
writer
|
type[DatasetWriter]
|
A concrete :class: |
required |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
Examples: