Would you be open to a small standalone tutorial showing how DVC Experiments can compare two image-augmentation policies for the same segmentation task, including the exact derived training data, representative image/mask previews, and the resulting Dice score?
User problem
Changing an augmentation policy changes the effective training dataset. A metric alone does not show which labeled samples the model actually saw, and an image preview alone does not preserve the dataset and policy that produced the result.
The current example-get-started-experiments already demonstrates the important DVC pieces for the swimming-pool segmentation task: a staged pipeline, parameters, cached outputs, experiment queues, Dice evaluation, and DVCLive image plots. Issue #188 also established that the test set must remain fixed for metrics and prediction images to be comparable across experiments.
The proposed tutorial would keep that test set and evaluation unchanged. It would vary only a versioned training-data preparation stage.
Proposed workflow
- Add an optional
augment stage after data_split. It reads each training image/mask pair and writes a materialized derived dataset.
- Select a bounded
light or strong policy through DVC parameters.
- Use AlbumentationsX to apply each sampled geometric transform to the image and mask together. A stable seed derived from the sample ID and variant index is passed as
invocation_seed, so rerunning the stage produces the same files independently of processing order.
- Validate that every output image/mask pair has matching spatial dimensions and that each mask still contains only the declared segmentation labels.
- Save the resolved policy from
policy.to_dict() next to the derived dataset and log a small fixed set of original/augmented image-and-mask previews with DVCLive.
- Train and evaluate the same model for both policies. DVC then connects the selected parameters, stage dependencies, cached dataset output, previews, and Dice metric for each experiment.
A compact stage boundary could look like this:
stages:
augment:
cmd: python src/augment.py
deps:
- data/train_data
- src/augment.py
params:
- base.random_seed
- augmentation
outs:
- data/train_augmented
- results/augment
The AX call inside that stage remains explicit:
augmented = policy(
image=image,
mask=mask,
invocation_seed=sample_seed,
)
The user can then answer two related questions from one experiment comparison: what did this policy do to the labeled training data, and how did that dataset change affect the fixed evaluation?
Scope and placement
This would be a tutorial/example contribution. It requires no DVC or DVCLive API change.
The augmentation stage is independent of the training framework, so the proposal does not depend on the current fastai implementation and can coexist with the migration discussed in #255.
Would maintainers prefer this as:
- an optional advanced variant of
example-get-started-experiments;
- a separate generated example repository; or
- a focused dvc.org tutorial that leaves the foundational example unchanged?
If the workflow fits the project, I can prepare the tutorial after your guidance on the surface and training-framework target.
Optional dependency boundary
The example would install AlbumentationsX only in its own environment. The current public package is AGPL-3.0-only and requires Python 3.10 or newer. Users install the PyTorch build appropriate for CPU, CUDA, or MPS before AlbumentationsX; PyTorch is intentionally not selected through the AlbumentationsX package metadata.
Would you be open to a small standalone tutorial showing how DVC Experiments can compare two image-augmentation policies for the same segmentation task, including the exact derived training data, representative image/mask previews, and the resulting Dice score?
User problem
Changing an augmentation policy changes the effective training dataset. A metric alone does not show which labeled samples the model actually saw, and an image preview alone does not preserve the dataset and policy that produced the result.
The current
example-get-started-experimentsalready demonstrates the important DVC pieces for the swimming-pool segmentation task: a staged pipeline, parameters, cached outputs, experiment queues, Dice evaluation, and DVCLive image plots. Issue #188 also established that the test set must remain fixed for metrics and prediction images to be comparable across experiments.The proposed tutorial would keep that test set and evaluation unchanged. It would vary only a versioned training-data preparation stage.
Proposed workflow
augmentstage afterdata_split. It reads each training image/mask pair and writes a materialized derived dataset.lightorstrongpolicy through DVC parameters.invocation_seed, so rerunning the stage produces the same files independently of processing order.policy.to_dict()next to the derived dataset and log a small fixed set of original/augmented image-and-mask previews with DVCLive.A compact stage boundary could look like this:
The AX call inside that stage remains explicit:
The user can then answer two related questions from one experiment comparison: what did this policy do to the labeled training data, and how did that dataset change affect the fixed evaluation?
Scope and placement
This would be a tutorial/example contribution. It requires no DVC or DVCLive API change.
The augmentation stage is independent of the training framework, so the proposal does not depend on the current fastai implementation and can coexist with the migration discussed in #255.
Would maintainers prefer this as:
example-get-started-experiments;If the workflow fits the project, I can prepare the tutorial after your guidance on the surface and training-framework target.
Optional dependency boundary
The example would install AlbumentationsX only in its own environment. The current public package is AGPL-3.0-only and requires Python 3.10 or newer. Users install the PyTorch build appropriate for CPU, CUDA, or MPS before AlbumentationsX; PyTorch is intentionally not selected through the AlbumentationsX package metadata.