A ManiSkill-based framework for studying OOD generalization and robustness of visual robot manipulation policies under environment distribution shifts.
This project investigates how observation design and domain randomization affect PPO policies for robotic pushing when object positions, goal distances, robot initial states, and physical parameters differ from the training distribution.
How well do visual robot manipulation policies generalize beyond their training distribution, and can domain randomization improve their robustness?
We build a GPU-parallel training and OOD evaluation framework based on ManiSkill + PPO, including:
- Up to 2048 parallel simulation environments
- State / Depth / Depth + Goal observation settings
- Modular Observation–Encoder architecture
- Episode-level Domain Randomization
- Multiple OOD evaluation conditions
- Multi-seed evaluation
- Rollout and trajectory-based failure analysis
We compare three observation settings:
| Policy | Observation |
|---|---|
| State | Privileged environment state |
| Depth | Depth image + robot proprioception |
| Depth + Goal | Depth image + task goal + robot proprioception |
For the PushCube task, the planar target region is not distinguishable from the tabletop in the depth image.
Therefore, the final visual policy uses:
Depth + Task Goal + Proprioception
where:
- Depth provides scene geometry and cube information
- Task Goal provides the desired target position
- Proprioception includes robot joint states and TCP pose
The cube position is not directly provided to the visual policy and must still be inferred from visual observation.
The policy is trained with PPO in GPU-parallel ManiSkill environments.
Main setup:
- Simulator:
physx_cuda - Parallel environments:
2048 - Observation:
depth - Encoder:
depth_goal - Total training steps:
2M - Training seeds:
1 / 2 / 3
Episode DR randomizes three factors during training:
| Factor | Randomization |
|---|---|
| Cube position | Expanded XY initialization range |
| Goal distance | Randomized pushing distance |
| Robot initial state | Increased joint-state initialization noise |
The network architecture, PPO hyperparameters, observation design, and depth preprocessing are kept fixed between the baseline and DR experiments.
Policies are evaluated under multiple held-out distribution shifts.
Cube OOD— unseen object initialization regionGoal Near— shorter pushing distanceGoal Far— longer pushing distanceQpos Shift— increased robot joint-state perturbationEpisode Combined— simultaneous cube, goal, and robot-state shifts
- Cube mass shift
- Contact friction shift
- Combined physics shift
Full Combined— simultaneous episode-level and physics distribution shifts
The primary evaluation metric is:
Success Rate at Episode End
Each policy is evaluated using 100 episodes per condition.
We compare the Depth + Goal baseline with Depth + Goal + Episode DR across three independently trained seeds.
| Condition | Depth + Goal | + Episode DR |
|---|---|---|
| Normal | 61.7 ± 20.2% | 70.3 ± 3.1% |
| Cube OOD | 29.7 ± 7.2% | 35.7 ± 4.0% |
| Goal Near | 93.7 ± 0.6% | 91.0 ± 1.7% |
| Goal Far | 8.0 ± 7.2% | 20.3 ± 12.9% |
| Qpos Shift | 64.3 ± 12.6% | 77.0 ± 3.6% |
| Episode Combined | 9.7 ± 5.5% | 16.0 ± 2.6% |
| Full Combined | 15.3 ± 9.1% | 19.7 ± 1.2% |
Values are mean ± standard deviation over three training seeds.
Each trained policy is evaluated on 100 episodes per condition using the same evaluation seed.
Episode DR improves most OOD conditions — up to +12.7 pp on Qpos Shift and +12.3 pp on Goal Far — while slightly degrading Goal Near (−2.7 pp), showing that the effect is not uniform across distribution shifts.
On the Normal condition, Episode DR reduces seed-level variance from ±20.2 pp to ±3.1 pp, substantially improving training stability.
Episode-level Domain Randomization improves both average performance and training stability in several conditions.
Notable observations include:
- Normal success improves from 61.7% to 70.3%
- Goal-distance OOD improves from 8.0% to 20.3%
- Qpos-shift success improves from 64.3% to 77.0%
- Training-seed variance is substantially reduced under several conditions
Cube OOD,Goal Far, and combined distribution shifts remain challenging
Physics shifts cause relatively limited degradation compared with episode-level distribution shifts, suggesting that the current task is more sensitive to spatial and initial-state shifts than to the tested mass/friction changes.
To understand policy failures beyond aggregate success rates, we analyze rollout videos together with privileged diagnostic trajectories.
The privileged information is used only for analysis, not as policy input.
We track:
- TCP-to-cube distance
- Cube-to-goal distance
- TCP-to-goal distance
- Cube and TCP trajectories
Failure cases suggest that OOD errors are not purely caused by object localization.
In many failed rollouts, the policy can approach the cube but fails to:
- maintain stable contact,
- select an effective pushing direction,
- sustain pushing for sufficiently long trajectories,
- recover when multiple distribution shifts occur simultaneously.
Goal Far remains particularly challenging, indicating limited generalization to long-horizon pushing behavior.
maniskill-robust-manipulation/
│
├── configs/ # Training configurations
│
├── src/
│ ├── envs/ # ManiSkill environments and OOD variants
│ ├── models/ # Observation encoders and PPO agent
│ ├── train.py # PPO training
│ └── evaluate.py # Unified OOD evaluation
│
├── tests/ # Environment / observation smoke tests
├── experiments/ # Experiment and analysis scripts
├── plot_figures.py # Generates all figures in results/figures/
│
├── results/
│ ├── figures/ # Main visualization results
│ └── *.csv # Evaluation results
│
├── videos/ # Rollout videos + diagnostic trajectories (distances.csv)
├── docs/ # Experiment notes and benchmark design
│
├── requirements.txt
└── README.md
Create a Python environment and install the required dependencies.
pip install -r requirements.txtCore dependencies include:
- ManiSkill
- PyTorch
- Gymnasium
- NumPy
- Pandas
- PyYAML
- Matplotlib
- TensorBoard
A CUDA-enabled GPU is recommended for large-scale parallel training.
Example: train the Depth + Goal + Episode DR policy.
python -m src.train \
--config configs/depth_goal_dr_gpu.yamlExample configuration:
env_id: PushCubeDepthGoalDR-v1
obs_mode: depth
encoder_type: depth_goal
sim_backend: physx_cuda
cuda: true
num_envs: 2048
num_steps: 20
total_timesteps: 2000000Training checkpoints are saved under:
runs/<experiment_name>/
Large checkpoints and TensorBoard logs are not included in the repository.
Example: evaluate one trained Depth + Goal + Episode DR policy on all OOD conditions.
python -m src.evaluate \
--checkpoint runs/depth_goal_dr_gpu_seed1/final_ckpt.pt \
--model-name depth_goal_dr_seed1 \
--case all \
--num-episodes 100 \
--seed 1000 \
--sim-backend physx_cuda \
--device cuda \
--obs-mode depth \
--encoder-type depth_goal \
--output results/depth_goal_dr_seed1_all.csvFor multi-seed experiments, all trained policies are evaluated using the same evaluation settings.
All figures in results/figures/ are generated by one script:
python plot_figures.py| Figure | Description |
|---|---|
ood_success_comparison.png |
Main grouped bar chart (mean ± std over 3 seeds) |
ood_improvement.png |
ΔSR = SR_DR − SR_Baseline per OOD condition |
seed_stability.png |
Seed-level dot plot on the Normal condition |
failure_analysis.png |
Observation design, failure rollout, distance diagnostics |
- GPU-parallel PPO training pipeline
- Unified OOD evaluation framework
- State / Depth / Depth + Goal observation framework
- Episode-level Domain Randomization
- Episode / Physics / Full OOD benchmark
- Three-seed baseline evaluation
- Three-seed Episode DR evaluation
- Rollout-based failure analysis
- TCP–cube–goal trajectory diagnostics
- Publication-quality result figures
- Factor-wise ablation of Cube / Goal / Qpos randomization
- Analysis of combined OOD failure modes
- Improved robustness under long-distance pushing
- Additional visual manipulation tasks
This repository focuses on understanding visual robot manipulation robustness under distribution shift, rather than maximizing performance on a single PushCube benchmark.
The current experiments aim to separate the effects of:
- observation design,
- visual partial observability,
- training distribution coverage,
- domain randomization,
- policy failure under OOD conditions.



