Skip to content

Latest commit

 

History

History
237 lines (188 loc) · 6.31 KB

File metadata and controls

237 lines (188 loc) · 6.31 KB

Workspace Structure

How Agentomics-ML organizes files during and after execution.

Directory Overview

agentomics-ml/
├── datasets/                 # Raw input datasets
├── prepared_datasets/        # Prepared public train/validation data
├── prepared_test_sets/       # Prepared hidden test data
├── workspace/                # Active execution workspace
│   ├── run/                  # Current run files
│   ├── best_iteration_snapshot/ # Best iteration snapshot
│   ├── reports/              # Iteration reports
│   └── extras/               # Logs and extra artifacts
└── outputs/                  # Final results

datasets/

Raw datasets use split folders:

datasets/my_dataset/
├── train/
│   ├── input/
│   ├── extras/             # Optional: supplementary training files
│   └── labels.csv
├── validation/             # Optional
│   ├── input/
│   ├── extras/             # Optional: supplementary training files
│   └── labels.csv
├── test/                   # Optional hidden test set
│   ├── input/
│   └── labels.csv
├── supplementary/          # Optional: supporting/supplementary materials
├── metadata.json           # Optional if task type is supplied at preparation
└── dataset_description.md  # Optional domain information

Each labels.csv must include id and numeric_label columns. Only train, validation, and test are supported split names. The input/ structure is recorded at preparation time and must match across all splits.

prepared_datasets/

After preparation, public splits are formatted for the agent:

prepared_datasets/my_dataset/
├── train/
│   ├── input/
│   ├── extras/             # If provided
│   └── labels.csv
├── validation/
│   ├── input/
│   ├── extras/             # If provided
│   └── labels.csv
├── supplementary/          # If provided
├── dataset_description.md
└── metadata.json

prepared_test_sets/

Test data is separated to keep it hidden:

prepared_test_sets/my_dataset/
└── test/
    ├── input/
    └── labels.csv

The agent never sees files in this directory during training.

workspace/

Active execution area:

workspace/run/

Current run working directory:

workspace/run/
├── shared/
│   ├── .conda/                  # Shared Conda environment
│   ├── config.json
│   ├── environment.yml
│   └── datasets/
├── current_iteration/
│   ├── current_step/            # Active step workspace
│   └── runtime_info/
├── iteration_0/                 # Archived iteration
├── iteration_1/
└── ...

workspace/best_iteration_snapshot/

Best iteration snapshot:

workspace/best_iteration_snapshot/
├── model_training/
│   ├── train.py
│   └── training_artifacts/
├── model_inference/
│   └── inference.py
├── runtime_info/
├── environment.yml
└── .conda/

Updated whenever a new best iteration is achieved.

workspace/run/shared/splits/

Versioned train/validation split folders:

workspace/run/shared/splits/
└── split_0/
    ├── train/
    │   ├── input/
    │   ├── extras/         # Optional
    │   └── labels.csv
    └── validation/
        ├── input/
        ├── extras/         # Optional
        └── labels.csv

Each time the agent changes the train/validation split, a new split_<n>/ folder is created. Iteration outputs record which split version they used. The input/ structure must match the original recorded structure across all splits. The extras/ subfolder may be created or modified by the agent.

workspace/reports/

Iteration reports are written here during runs. These are copied to outputs/<agent_id>/reports/ after completion.

workspace/extras/

Logs and auxiliary artifacts are stored here and copied to outputs/<agent_id>/extras/.

outputs/

Final results after run completion:

outputs/<agent_id>/
├── best_iteration_snapshot/           # Best iteration artifacts
│   ├── model_training/
│   │   ├── train.py
│   │   └── training_artifacts/
│   ├── model_inference/
│   │   └── inference.py
│   ├── runtime_info/
│   ├── environment.yml
│   └── .conda/
├── run/                      # All iterations + data splits
│   ├── shared/
│   │   ├── config.json
│   │   └── splits/
│   │       └── split_0/
│   │           ├── train/
│   │           │   ├── input/
│   │           │   └── labels.csv
│   │           └── validation/
│   │               ├── input/
│   │               └── labels.csv
│   ├── iteration_0/
│   ├── iteration_1/
│   └── ...
├── reports/
│   ├── markdown/
│   │   ├── run_report_iter_0.md
│   │   ├── run_report_iter_1.md
│   │   └── ...
│   └── pdf/
│       ├── iteration_0.pdf
│       ├── iteration_1.pdf
│       └── plots/
├── extras/                   # Additional files and logs
└── README.md                 # Run summary

File Notes

Iteration contents and artifact names can vary by run. Use <step_id>/output.json inside each archived iteration or best iteration snapshot as the structured source of truth for step outputs. Use outputs/<agent_id>/README.md for the most accurate per-run details.

Cleanup

Remove Specific Run

rm -rf outputs/<agent_id>

Clean Workspace

rm -rf workspace/run/*
rm -rf workspace/best_iteration_snapshot/*

Clean Everything

rm -rf outputs/*
rm -rf workspace/*
rm -rf prepared_datasets/*
rm -rf prepared_test_sets/*

Docker Volumes

In Docker mode, workspace is mounted as a volume:

  • Code repository: read-only
  • Workspace: read-write
  • Outputs: read-write

This isolates agent execution from the host system.

Related