Cloud inference and training for SmolVLA (Small Vision-Language-Action) models with Cyberwave robot control integration.
This cloud node runs SmolVLA inference (deploy.py + CwProcessor) and optional training (train.py + CwTrainer) on Cyberwave infrastructure, enabling real-time robot control from language instructions and camera observations.
┌─────────────────────────────────────────────────────────────────────┐
│ Cyberwave Cloud Node │
│ ┌───────────────────────────────────────────────────────────────┐ │
│ │ deploy.py │ │
│ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────────┐ │ │
│ │ │ SmolVLA │───▶│ predict_fn │───▶│ Action Chunk │ │ │
│ │ │ Policy │ │ │ │ (50 actions) │ │ │
│ │ └─────────────┘ └─────────────┘ └─────────────────┘ │ │
│ └───────────────────────────────────────────────────────────────┘ │
│ │ │
│ ┌───────────────────────────▼───────────────────────────────────┐ │
│ │ CwProcessor │ │
│ │ • Cyberwave SDK client (auto-configured) │ │
│ │ • Background camera fetchers (daemon threads) │ │
│ │ • MQTT subscription (joint states) │ │
│ │ • Action publishing (MQTT) │ │
│ └───────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
│
▼
┌──────────────────┐
│ Robot + Cameras │
│ (via Cyberwave) │
└──────────────────┘
- Weights Download: Model weights are fetched from Cyberwave MLModel API (signed URLs) and cached locally
- Model Loading: SmolVLA policy is loaded from the downloaded checkpoint
- Camera Binding: Background daemon threads continuously fetch frames from each camera twin
- Observation Collection:
CwProcessorreads cached frames and joint states - Inference: The model predicts a chunk of 50 actions from images + state + instruction
- Execution: Actions are published to the robot via MQTT
- Loop: Process repeats for
max_stepsiterations
Each camera twin has a dedicated daemon thread that continuously polls for frames:
┌─────────────────────────────────────────────────────────────┐
│ Camera Threads (Background) │
│ │
│ camera_wrist thread ──► GET /twins/{uuid}/latest-frame │
│ │ │ │
│ └──────────────────────▶ cache (np.ndarray + bytes) │
│ │
│ camera_front thread ──► GET /twins/{uuid}/latest-frame │
│ │ │ │
│ └──────────────────────▶ cache (np.ndarray + bytes) │
│ │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Control Loop │
│ │
│ get_inputs() ──► reads cached frames (instant, no I/O) │
│ │
└─────────────────────────────────────────────────────────────┘
This decouples frame fetching from the inference loop, ensuring consistent frame rates.
Training configs contain camera names (e.g., camera_wrist, camera_front). At runtime, cameras are mapped from camera_endpoints_by_role:
camera_endpoints_by_role: {
"camera_wrist": "https://api.cyberwave.com/api/v1/twins/{uuid}/latest-frame",
"camera_front": "https://api.cyberwave.com/api/v1/twins/{uuid}/latest-frame"
}
↓
Extract UUIDs from URLs
↓
Fetch Twin objects via Cyberwave SDK
↓
Start background fetcher threads
# Create virtual environment
python3 -m venv ~/.venv/smolvla
source ~/.venv/smolvla/bin/activate
# Install dependencies
./install.shThe install script uses requirements.txt which includes:
numpy,Pillow,requests- Core dependencieslerobot[smolvla]- LeRobot with SmolVLA supportpeft- LoRA fine-tuningcyberwave>=0.3.46- Cyberwave SDKzstandard- For.tar.zstweight archives
| Variable | Required | Description |
|---|---|---|
CYBERWAVE_API_KEY |
Yes | Cyberwave API key (set by Cloud Node) |
SMOLVLA_CHECKPOINT |
No | Override checkpoint path (otherwise uses weights_url) |
The Cloud Node sends a JSON payload with robot and camera configuration:
{
"robot_twin_uuid": "uuid-of-robot-twin",
"instruction": "pick up the red block and place it in the bin",
"weights_url": "https://api.cyberwave.com/api/v1/mlmodels/{uuid}/weights",
"policy_repo_id": "lerobot/smolvla_base",
"camera_endpoints_by_role": {
"camera_wrist": "https://api.cyberwave.com/api/v1/twins/{uuid}/latest-frame",
"camera_front": "https://api.cyberwave.com/api/v1/twins/{uuid}/latest-frame"
},
"twin_calibration": {
"leader": { "...": "..." },
"follower": { "...": "..." }
},
"calibration_robot_type": "follower",
"mode": "live",
"max_steps": 1000,
"wait_for_joint_update_seconds": 1.0,
"actions_per_cycle": 25,
"action_sleep_seconds": 0.1,
"inference_loop": true
}The weights_url points to the Cyberwave MLModel API endpoint which returns a signed URL:
GET /api/v1/mlmodels/{uuid}/weights
→ { "signed_url": "https://storage.../checkpoint.tar.zst", "expires_at": "..." }
GET {signed_url}
→ Download and extract .tar.zst archive
Supported archive formats:
.tar.zst(Zstandard compressed tar).tar.gz/.tgz(Gzip compressed tar).zip(ZIP archive)
Weights are cached at ~/.cache/cyberwave/weights/ with automatic directory resolution to find config.json.
| Parameter | Default | Description |
|---|---|---|
mode |
— | e.g. live (passed through from platform; informational) |
max_steps |
1 | Maximum total actions to execute |
wait_for_joint_update_seconds |
1.0 | Timeout waiting for first joint state after MQTT subscribe |
actions_per_cycle |
25 | Actions to execute per inference cycle (from 50-action chunk) |
action_sleep_seconds |
0.1 | Sleep time between publishing each action |
inference_loop |
true | If true, run multiple inference cycles; if false, single chunk |
camera_poll_interval_seconds |
0.05 | Background camera polling interval |
export CYBERWAVE_API_KEY=your-api-key
python deploy.py /path/to/params.jsonThe cyberwave.yml configures the Cloud Node:
cyberwave-cloud-node:
install_script: ./install.sh
inference: |
source "$HOME/.venv/smolvla/bin/activate" && \
python deploy.py {body}
profile_slug: smolvlasmolvla/
├── deploy.py # Inference entry point (model load + CwProcessor)
├── train.py # Training entry point (CwTrainer + lerobot)
├── cw_processor.py # Inference: SDK, MQTT, cameras, weights download
├── cw_trainer.py # Training: dataset download, metrics, artifact packaging
├── smolvla_resolver.py # Model-specific metadata (camera mapping, config parsing)
├── smolvla_trainer.py # SmolVLA TrainPipelineConfig builder
├── base_resolver.py # Abstract resolver interface
├── base_trainer.py # Abstract trainer interface
├── requirements.txt # Python dependencies
├── install.sh # Installation script
├── cyberwave.yml # Cloud Node configuration
├── ARCHITECTURE.md # Detailed architecture documentation
└── README.md
- Parses request payload and downloads weights via
download_weights() - Loads SmolVLA policy from checkpoint
- Builds
predict_fnfor inference - Creates
CwProcessorand runs the control loop
download_weights(): Fetches weights from MLModel API, extracts archives, caches locallyCwProcessor: Orchestrates all Cyberwave interactions- Creates SDK client with auto-configured MQTT
- Spawns background camera fetcher threads
- Subscribes to joint state updates
- Publishes action predictions via MQTT
InferenceRequest: Dataclass for request parametersCameraBinding: Holds cached frame data per camera
- Loads
train_config.jsonfrom checkpoint - Extracts training camera names and dimensions
- Builds camera mapping (training name → runtime key)
Terminal output shows real-time progress:
══════════════════════════════════════════════════
CYBERWAVE SETUP
══════════════════════════════════════════════════
API Key: cw_your_api_key
✓ Client created
✓ MQTT connected
✓ Joints received: [-0.01, 0.01, 0.07, 0.02, -0.00, 0.12]
✓ Camera camera_wrist ready
✓ Camera camera_front ready
Started 2 background camera fetchers
══════════════════════════════════════════════════
SMOLVLA CONTROL LOOP
══════════════════════════════════════════════════
Max steps: 1000
Actions/cycle: 25
──────────────────────────────────────────────────
CYCLE 1 │ 0/1000 steps (0%)
──────────────────────────────────────────────────
Cameras: 2 frames │ State dim: 6
Running inference...
Predicted 50 actions -> executing 25
[████████████████████] 25/25 ✊ [+0.11, -0.00, +0.19, +0.41, -0.15, +0.10]
✓ Executed 25 actions
JSON result is written to stdout for orchestration:
{
"status": "ok",
"robot_twin_uuid": "...",
"instruction": "put object in box",
"steps_executed": 1000,
"initial_joints": {"_1": -0.01, "_2": 0.01, ...}
}- Python >= 3.10
- CUDA-capable GPU (recommended)
- LeRobot with SmolVLA support
- Cyberwave SDK >= 0.3.46
- LeRobot - Robot learning framework
- SmolVLA - Small Vision-Language-Action model
- Cyberwave - Robot cloud infrastructure
See LICENSE file for details.