Available on Zenodo
What is contained?
- We provide the code and the scripts to train policies with BC, PPO, OBC, OBC-PPO, RL fine-tuning and self-distillation. We thus also provide expert policies for imitation learning.
- We also include pretrained checkpoints for OBC and OBC-PPO, to reproduce results and videos.
To reproduce the training experiments, we are providing you two ways to set up a feasible environment.
The installation should usually take less than 30 minutes on a modern computer with fast internet connection.
Both methods install the same package set, defined once in environment.yml (conda) and docker-cuda/requirements.txt (Docker).
Use the provided Dockerfile, you can create a docker container that can be then used to run all the arnold experiments (note: this assume that Docker is installed in your system).
To build the Docker image, navigate to the docker-cuda directory containing the Dockerfile and run:
docker build -t arnold_image .The image contains only the Python environment. The repository itself is mounted into
the container at run time, so that data/ (which you download separately from Zenodo)
and everything the runs write to output/ live on the host and survive the container.
From the repository root:
docker run -it --rm --gpus all \
-v "$PWD":/arnold -w /arnold \
arnold_image /bin/bashThis will start an interactive session within the container, where you can then execute the training or evaluation scripts exactly as written in the sections below, for example:
python src/main_bc_ppo.py --task elbow_pose --num_envs 16 --num_steps 5000000 --localNotes:
- Drop
--gpus allif the host has no NVIDIA GPU (this also requires the NVIDIA Container Toolkit). The image still runs CPU-only, in which case pass--device cputo the scripts. On Apple Silicon the image builds and runs natively, but CPU-only. - The image renders headlessly through OSMesa (
MUJOCO_GL=osmesa), so--renderworks without a display. With a GPU attached,-e MUJOCO_GL=eglis faster. - On Linux, files the container writes into the mounted repository are owned by
root. Add-u "$(id -u):$(id -g)"to thedocker runcommand to keep them owned by you.
The quickest way is to create the environment from the provided file:
conda env create -f environment.yml
conda activate arnold
pip install imitation==1.0.0The trailing pip install imitation is required and must come last. MyoSuite 2.2.0 pins
gym==0.13, whose metadata demands cloudpickle~=1.2.0, while imitation==1.0.0 pulls
in huggingface-sb3, which demands cloudpickle>=1.6. pip cannot satisfy both in a
single resolution, so imitation is installed in a second pass; this upgrades
cloudpickle to 3.x, which is the version every package actually runs against. Only
gym's stale pin — for a code path this project never uses — is left unsatisfied, and
pip check will report it.
Equivalently, you can create the environment and install the dependencies manually:
conda create -n arnold python=3.8
conda activate arnold
pip install \
cloudpickle==1.2.2\
gym==0.13.0\
gymnasium==0.29.1\
h5py==3.7.0\
wandb\
tqdm\
numpy\
ipdb
pip install stable-baselines3==2.2.1
pip install MyoSuite==2.2.0
pip install sb3-contrib==2.2.1
pip install Shimmy==1.3.0
pip install imageio
# Needed by the analysis and plotting scripts
pip install \
matplotlib\
seaborn\
scikit-learn\
scipy\
pandas\
tensorboard\
joblib
# Must come last, for the reason explained above
pip install imitation==1.0.0You may need to install some opengl-related system packages:
apt-get update && apt-get install -y libgl1-mesa-glx libosmesa6The git repository contains code and environment/expert configuration files. Benchmark results, model weights, training logs and analysis data are external inputs. The release record is Zenodo. See this record for file structure.
Environment and expert configurations are already included under data/env_configs/ and
data/expert_configs/ without need to download from Zenodo. External inputs from Zenodo use this layout:
| Directory | Contents | Needed for |
|---|---|---|
data/final_benchmarks/ |
Per-method/per-seed result JSONs and expert references; arnold_single_task/ and transfer_learning/ logs; example_training_curve/ base OBC logs; example_checkpoint/ with one checkpoint and matching normalization file. |
Performance, ablation, student, transfer and fine-tuning plots; checkpoint-loading examples. |
data/final_benchmarks_extra/ |
CSI evaluation results and csi*/training/ logs; bilateral and normalization ablations; mt-curves/ caches; rl_finetuning/ logs. |
CSI, baseline and RL fine-tuning plots. |
data/analysis/ |
Curve caches, reproduction/ metadata, compact hand signals, human/simulation EMG inputs, gait recordings and exported tables. |
Learning curves, PCA/NMF summaries, hand and EMG analyses. |
data/final_checkpoints/ |
Separately released model checkpoints and associated normalization/configuration files. | Evaluation, fresh recordings and resumed training beyond the bundled example. |
data/expert_policies/ |
Expert policy checkpoints (EXPERT_POLICIES_PATH). |
Expert evaluation and rollout collection. |
data/kinesis/ |
Locomotion model assets. | Instantiating the Kinesis environment. |
Extract the packages from the repository root:
mkdir -p data
tar -xzf final-benchmarks.tar.gz -C data
tar -xzf final-benchmarks-extra.tar.gz -C data
tar -xzf expert-policies.tar.gz -C data
tar -xzf analysis.tar.gz -C dataSee Data and checkpoints for input requirements and cache behavior.
For the Task-SV sensorimotor vocabulary ablation, see training and evaluation instructions.
Run the training stages in order:
- Train OBC from scratch for 50M steps.
- Continue OBC for 5M steps at the reduced learning rate.
- Train the final Arnold agent from the 55M-step checkpoint using super-expert policies.
This command starts an On-policy Behavioral Cloning (OBC) training from scratch using all the 14 tasks.
It assumes expert demonstrations are available when imitation_coef > 0.
python src/main_bc_ppo_multi_task.py \
--tasks hand_thumb_reach hand_index_reach hand_middle_reach hand_ring_reach hand_little_reach \
reorient pen baoding_p1_cw baoding_p1_ccw baoding_p2 baoding_p2_overlap elbow_pose relocate kinesis kinesis \
relocate baoding_p1_cw baoding_p2 baoding_p2_overlap kinesis kinesis \
relocate baoding_p1_ccw baoding_p2 baoding_p2_overlap kinesis kinesis \
--num_envs_per_task 2 \
--ent_coef=0 \
--vf_coef=0.5 \
--pg_coef=0 \
--imitation_coef=1 \
--num_steps=50000000 \
--batch_size=128 \
--rollout_steps=512 \
--embedding_size=128 \
--dim_feedforward=512 \
--num_heads=4 \
--num_layers=6 \
--lr=1e-3 \
--log_interval=1 \
--n_epochs=3 \
--separate_vf_decoder \
--policy_outputs_variance \
--norm_reward \
--dense_reward \
--out_prefix=obc_ \
--seed 1 \
--project_name arnold_obc_multi_task_scratchAfter the first 50M steps, OBC (as well as BC and OBC-PPO) is continued for 5M additional steps at a smaller learning rate, which boosts the performance in certain tasks. This is the same command as above with --load_path set to the 50M checkpoint, --num_steps=5000000 and --lr=1e-5. The resulting 55M-step checkpoint (rl_model_54974700_steps.zip) is the starting point for the RL fine-tuning and for the final Arnold agent.
This command trains the final Arnold agent by imitating specified super-expert policies, resuming from the 55M-step OBC model.
python src/main_bc_ppo_multi_task.py \
--tasks hand_thumb_reach hand_index_reach hand_middle_reach hand_ring_reach hand_little_reach \
reorient pen baoding_p1_cw baoding_p1_ccw baoding_p2 baoding_p2_overlap elbow_pose relocate kinesis kinesis \
relocate baoding_p1_cw baoding_p2 baoding_p2_overlap kinesis kinesis \
relocate baoding_p1_ccw baoding_p2 baoding_p2_overlap kinesis kinesis \
--load_path data/final_checkpoints/obc/seed_0 \
--num_envs_per_task 2 \
--ent_coef=0 \
--vf_coef=0.5 \
--pg_coef=0 \
--imitation_coef=1 \
--num_steps=10000000 \
--batch_size=128 \
--rollout_steps=512 \
--embedding_size=128 \
--dim_feedforward=512 \
--num_heads=4 \
--num_layers=6 \
--lr=1e-5 \
--log_interval=1 \
--n_epochs=3 \
--separate_vf_decoder \
--policy_outputs_variance \
--norm_reward \
--custom_experts data/expert_configs/arnold_experts_seed_1.json \
--dense_reward \
--out_prefix=285_ \
--seed 1 \
--project_name arnold_final_bc_super_expertsHere's how to configure src/main_bc_ppo_multi_task.py for these and other training types in more detail:
- PPO: Set
--imitation_coef=0,--pg_coef=1,--ent_coef=1e-6, and--lr=2e-5. - OBC-PPO: Similar to PPO, but set
--imitation_coef=1and--lr=1e-3(reduced to1e-5for the additional 5M steps). - BC: This involves training with
imitation_coef > 0andpg_coef = 0. The 'Final Arnold agent' command above is an example. If performing online imitation of an expert policy, the--use_expert_actionsflag is typically used. - Super-experts: Same as PPO, but resume from the 55M-step OBC checkpoint. Use a low learning rate (
--lr=2e-6) and 32 instances of the same task (e.g.,--num_envs_per_task=32 --tasks <single_task_name>). The standard deviation of the action distribution approaches 0 during OBC pretraining and must be reset for the fine-tuning to improve over the expert: pass--reset_std --log_std_init -3.
The src/benchmark.py script allows you to evaluate the performance of various pretrained models, including those trained with OBC, Arnold, and expert policies.
To test a model trained with OBC or Arnold, you need to specify the path to the saved model (.zip file) and the task you want to evaluate. We provide trained models that you can download from Zenodo.
Here's an example command:
python src/benchmark.py \
--load path/to/your/model.zip \
--task <task_names> \
--arnold \
--num_episodes <number_of_episodes> \
--deterministic \
--device <cpu_or_cuda> \
--render- Replace
path/to/your/model.zipwith the actual path to your trained model. - Set
<task_names>to one or many of the available tasks. - Include the
--arnoldflag if the model was trained with Arnold. - Adjust
--num_episodesto set how many episodes to run for evaluation. The results reported in the paper use 200 episodes per task. - Use
--deterministicfor deterministic actions from the policy. - Specify the
--device(e.g.,cpuorcuda). - Optionally render to video with
--render(important: on Mac this requires runningmjpythoninstead ofpython)
Example for an Arnold model:
python src/benchmark.py \
--load data/final_benchmarks/example_checkpoint/rl_model_64670238_steps.zip \
--task kinesis \
--arnold \
--num_episodes 10 \
--deterministic \
--device cpuExample for a OBC model:
python src/benchmark.py \
--load data/final_checkpoints/obc/seed_0/rl_model_54974700_steps.zip \
--task relocate \
--arnold \
--num_episodes 10 \
--deterministic \
--device cpuUsually to reproduce these evaluations, the runtime is less than 1 minute.
To evaluate an expert policy, you need to specify the task.
python src/benchmark.py \
--task <task_name> \
--expert \
--num_episodes <number_of_episodes> \
--deterministic \
--device <cpu_or_cuda>- Set
<task_name>to one of the available tasks. - Include the
--expertflag to indicate you are testing an expert policy.
Example for an expert policy:
python src/benchmark.py \
--task relocate \
--expert \
--num_episodes 10 \
--deterministic \
--device cpu \
--renderYou can choose <task_name> from the following list:
baoding_p1_ccwbaoding_p1_cwbaoding_p2baoding_p2_overlaphand_thumb_reachhand_index_reachhand_middle_reachhand_ring_reachhand_little_reachpenrelocatereorientelbow_posekinesis
Scripts for performance plots and benchmark tables. Figures are written under data/figures/.
-
Script:
plotting/plot_radar.py -
Description: Radar chart comparing multi-task performance (PPO, BC, Arnold) against the single-task experts.
-
Output:
data/figures/radar_plot_ppo_bc_arnold.png/.svg. -
Example Usage:
python plotting/plot_radar.py
-
Script:
plotting/plot_ppo_ablation_bars.py -
Description: Per-task bar plot comparing PPO variants (PPO, w/o reward norm, w/o observation norm, MT-PPO) against the experts.
-
Output:
data/figures/bar_plot_ppo_ablation.png/.svg. -
Example Usage:
python plotting/plot_ppo_ablation_bars.py
-
Script:
plotting/plot_arnold_ablation_bars.py -
Description: Per-task bar plot comparing agent variants (BC, OBC, OBC-PPO, Arnold) against the experts, plus an improvement-over-baseline version.
-
Output:
data/figures/bar_plot_arnold_ablation.png/.svganddata/figures/bar_plot_arnold_ablation_improvement.png/.svg. -
Example Usage:
python plotting/plot_arnold_ablation_bars.py
python plotting/ablation_table.pyWrites CSV and LaTeX tables to data/analysis/tables/, reporting relative solved-step
fractions, SEM across seeds, and task-paired Wilcoxon tests with Holm correction.
python plotting/plot_relative_dotplot.pyCompares per-task solved-step fractions for obc_task_sv against obc by default.
Use --method and --reference to select policies. Writes data/figures/relative_dotplot.svg.
python plotting/plot_capacity_performance.pyCompares OBC model sizes against expert performance. Writes capacity.svg and
performance.csv under data/figures/capacity/. Error bars show episode SEM for
one run per model size.
python plotting/plot_bilateral_reward.pyReads data/final_benchmarks_extra/bilateral/bilateral.json and compares episode
rewards. Writes data/figures/bilateral_reward.svg.
See Performance plots and tables for input requirements and benchmark selections.
Principal Component Analysis (PCA) of the action space of trained policies, to study the effective dimensionality of the learned actions. To generate new inactivation results:
- Collect activations and actions from a trained policy.
- Fit PCA or NMF components and evaluate inactivation performance.
- Plot the resulting inactivation curves.
Cumulative explained variance uses the recordings from the first step and can be plotted independently of the inactivation analysis.
-
Script:
plotting/collect_activations.py -
Description: Rolls out a trained policy and saves per-episode observations, action means, rewards, and (for Arnold) intermediate activations. The action means feed the PCA steps below.
-
Output: HDF5 (
.h5) files underdata/activations/<policy_id>/(e.g.,data/activations/example_64670238/). -
Example Usage:
# Collect activations for multiple tasks for task in hand_thumb_reach hand_index_reach hand_middle_reach hand_ring_reach hand_little_reach reorient pen baoding_p1_ccw baoding_p1_cw baoding_p2 baoding_p2_overlap; do python plotting/collect_activations.py \ --load data/final_benchmarks/example_checkpoint/rl_model_64670238_steps.zip \ --task $task \ --num_episodes 100 \ --arnold \ --normalize \ --device cpu \ --out_dir data/activations done
Use the existing per-episode recordings or compact signals with explicit paths:
python plotting/analyze_pca_inactivation.py \
--load path/to/rl_model_64670238_steps.zip \
--signals data/activations/example_64670238 --method pca --scope global \
--out_dir data/pca_analysis/example_64670238/globalThe command writes curves.csv, episode results and fitted components. Use --scope task
for per-task fits. NMF requires physical controls from the compact signal collector.
See CSI analysis for the methods and cached-result provenance.
python plotting/plot_pca_inactivation.py \
--curves data/pca_analysis/example_64670238/global/curves.csvFigures are saved to data/figures/pca_inactivation/. Without --curves, the included
historical PCA/NMF summary is plotted.
-
Script:
plotting/plot_action_pca_variance.py -
Description: Plots the cumulative explained variance of the actions (per-task and global) versus the number of principal components.
-
Output:
data/figures/cumulative_variance/<policy_id>/(.png/.svg). -
Example Usage:
python plotting/plot_action_pca_variance.py --activations_dir data/activations/example_64670238 --out_dir data/figures/cumulative_variance/example_64670238
python plotting/analyze_subspaces.py --data_dir data/analysis/signals
python plotting/analyze_subspaces.py --selection capacity_recordings --successful 100Compares control subspaces using PVD and PAD and writes measurements and figures to
data/figures/subspaces/. See Control subspace comparisons
for policy selections and interpretation, and Analysis signals for
collecting or importing the required recordings.
python plotting/analyze_smoothness.py
python plotting/analyze_baoding_kinematics.pyUses compact hand recordings under data/analysis/signals/. Smoothness metrics and
figures go to data/figures/smoothness/; Baoding dimensionality and human-comparison
figures go to data/figures/baoding_pca/. See Hand smoothness and Baoding PCA
for inputs, metric definitions and recording commands.
python plotting/analyze_emg.py
python plotting/analyze_gait_factors.pyUses human and simulation profiles under data/analysis/emg/ and writes CSVs and SVGs
to data/figures/emg/ and data/figures/gait_factors/. See EMG and gait factors
for human-profile import, gait collection, segmentation and analysis methods.
The released benchmark results and training logs in data/final_benchmarks_extra/ already cover every arm, so to reproduce only the figures, skip to step 5. Steps 1-4 regenerate the experiments from scratch.
Below, <task> is one of the 14 tasks listed above and <dim> is one of the seven action-space sizes used in the paper: 1, 2, 5, 10, 20, 30, 40.
Run the experiment in this order:
- Train the base MLP policy with OBC.
- Extract its CSI action subspace.
- Fine-tune the OBC and PPO arms inside that subspace; keep the frozen arm untrained.
- Benchmark the frozen and fine-tuned arms.
- Generate the performance and learning-curve figures.
Trains the unconstrained MLP policy with On-policy Behavioral Cloning (imitation_coef=1, pg_coef=0).
python src/main_bc_ppo.py \
--task <task> \
--network csi \
--num_envs 16 \
--imitation_coef 1.0 \
--pg_coef 0 \
--vf_coef 0 \
--num_steps 5000000 \
--localThe checkpoint is written to output/training/ongoing/<run_name>/rl_model_5000000_steps.zip.
Rolls the trained policy out and runs PCA on its actions to obtain the control subspace.
python src/main_csi_get_subspace.py \
--policy_path output/training/ongoing/<run_name>/rl_model_5000000_steps.zip \
--task <task> \
--num_envs 4 \
--num_steps 100000 \
--save_path output/<task>_csiThis writes output/<task>_csi/subspace.npy and output/<task>_csi/mean.npy.
Three arms are compared, all constrained to the same subspace:
a. Frozen (black) — the policy is constrained but not trained further. Nothing to run here; it is evaluated directly from the step-1 checkpoint in step 4.
b. OBC fine-tuned (red) — 5M additional steps of On-policy Behavioral Cloning inside the subspace:
python src/main_bc_ppo.py \
--task <task> \
--network csi \
--num_envs 16 \
--imitation_coef 1.0 \
--pg_coef 0.0 \
--vf_coef 0.5 \
--ent_coef 0.0 \
--load_path output/training/ongoing/<run_name>/rl_model_5000000_steps.zip \
--load_csi_subspace output/<task>_csi/subspace.npy \
--csi_subspace <dim> \
--num_steps 5000000 \
--out_prefix 666_<task>_csi<dim>_c. PPO fine-tuned (blue) — 5M additional steps of PPO inside the subspace. This is the same script with the imitation term switched off (imitation_coef=0, pg_coef=1). --load_vecnormalize restores the observation-normalization statistics saved next to the checkpoint, which RL fine-tuning requires:
python src/main_bc_ppo.py \
--task <task> \
--network csi \
--num_envs 16 \
--imitation_coef 0 \
--pg_coef 1 \
--vf_coef 0.8 \
--ent_coef 1e-6 \
--load_path output/training/ongoing/<run_name>/rl_model_5000000_steps.zip \
--load_vecnormalize \
--load_csi_subspace output/<task>_csi/subspace.npy \
--csi_subspace <dim> \
--num_steps 5000000 \
--out_prefix 333_<task>_csi<dim>_Both fine-tuning arms resume from the 5M-step base checkpoint, so their own checkpoints are saved at rl_model_10000000_steps.zip.
Evaluate every (task, <dim>) pair and write the results where the plotting scripts look for them:
| Arm | Checkpoint to evaluate | --out_dir |
|---|---|---|
| Frozen (black) | step-1 base checkpoint (5M) | data/final_benchmarks_extra/csi_notrain_server/csi<dim>_all |
| OBC fine-tuned (red) | step-3b checkpoint (10M) | data/final_benchmarks_extra/csi_bc_server/csi<dim>_all |
| PPO fine-tuned (blue) | step-3c checkpoint (10M) | data/final_benchmarks_extra/csi_server/csi<dim>_all |
For the frozen arm the subspace is applied at evaluation time, so pass --csi_components:
python src/benchmark.py \
--policy csi \
--load output/training/ongoing/<run_name>/rl_model_5000000_steps.zip \
--task <task> \
--csi_components output/<task>_csi/subspace.npy \
--csi_subspace <dim> \
--normalize --deterministic \
--num_episodes 200 \
--device cpu \
--save_results \
--out_dir data/final_benchmarks_extra/csi_notrain_server/csi<dim>_allFor the fine-tuned arms the projection is already stored in the checkpoint, so --csi_components is omitted:
python src/benchmark.py \
--policy csi \
--load output/training/ongoing/<finetuned_run>/rl_model_10000000_steps.zip \
--task <task> \
--csi_subspace <dim> \
--normalize --deterministic \
--num_episodes 200 \
--device cpu \
--save_results \
--out_dir data/final_benchmarks_extra/csi_bc_server/csi<dim>_all(On macOS, rendering requires mjpython instead of python.)
-
Script:
plotting/plot_csi_analysis.pyandplotting/plot_csi_curves.py -
Description:
plot_csi_analysis.pyplots final performance of the frozen and fine-tuned policies as a function of action-space size;plot_csi_curves.pyplots the corresponding fine-tuning learning curves. -
Output:
data/figures/csi_analysis/(.png/.svg). -
Example Usage:
python plotting/plot_csi_analysis.py python plotting/plot_csi_curves.py
Both baselines are trained with the same script, src/main_sac_multi_task.py (multi-task SAC/PPO with an MLP policy); MT-PPO is selected with --algo ppo.
To generate new baseline benchmark results:
- Train the chosen baseline, MT-SAC or MT-PPO.
- Benchmark its saved checkpoints.
- Generate the performance figures from the benchmark results.
Learning curves use training histories and do not require benchmarking.
MT-SAC:
python src/main_sac_multi_task.py \
--project_name arnold-new-exp \
--tasks hand_thumb_reach hand_index_reach hand_middle_reach hand_ring_reach hand_little_reach \
reorient pen baoding_p1_ccw baoding_p1_cw baoding_p2 baoding_p2_overlap elbow_pose relocate kinesis \
--num_envs_per_task 2 \
--num_steps 50_000_000 \
--hidden_size 512 \
--num_layers 2 \
--train_freq 64 \
--save_freq 100000 \
--norm_reward \
--lr 1e-5 \
--out_prefix sac-run- \
--seed 771MT-PPO — the same script with --algo ppo and a rollout length:
python src/main_sac_multi_task.py \
--project_name arnold-mtppo-new \
--algo ppo \
--rollout_steps 256 \
--tasks hand_thumb_reach hand_index_reach hand_middle_reach hand_ring_reach hand_little_reach \
reorient pen baoding_p1_ccw baoding_p1_cw baoding_p2 baoding_p2_overlap elbow_pose relocate kinesis \
--num_envs_per_task 2 \
--num_steps 50_000_000 \
--hidden_size 512 \
--num_layers 2 \
--train_freq 64 \
--save_freq 100000 \
--norm_reward \
--lr 1e-5 \
--log_interval 1 \
--out_prefix sac-run- \
--seed 771Both baselines are evaluated with src/benchmark_multi_task_mlp.py. The task order, environment id and algorithm are recovered from the args.json saved next to each checkpoint, so only the checkpoint path is needed. Evaluate the latest checkpoint of each seed:
python src/benchmark_multi_task_mlp.py \
--load <run_dir>/rl_model_<steps>_steps.zip \
--num_episodes 200 \
--deterministic \
--device cpu \
--save_results \
--out_dir data/final_benchmarks/mt_ppo/seed_0- Learning curves:
plotting/plot_mt_algos.py(see Plotting Learning Curves). - MT-PPO bar:
plotting/plot_ppo_ablation_bars.pyaggregatesdata/final_benchmarks/mt_ppo/for the MT-PPO bar of the PPO ablation plot.
Learning-curve figures. Figures are written under data/figures/.
-
Inputs:
data/final_benchmarks_extra/rl_finetuning/anddata/final_benchmarks/example_training_curve/(TensorBoard logs). -
Script:
plotting/plot_rl_finetuning_curves.py -
Description: Compares the base multi-task OBC policy against several single-task policies fine-tuned with PPO, plotting the solved fraction versus training steps from TensorBoard logs. Experiment paths are set near the top of the script.
-
Output:
data/figures/rl_finetuning_combined/rl_finetuning_combined_solved_curves.png/.svg. -
Example Usage:
python plotting/plot_rl_finetuning_curves.py
-
Script:
plotting/plot_mt_algos.py -
Description: Plots multi-task RL baseline learning curves comparing MT-SAC and MT-PPO across all tasks.
-
Output:
data/figures/mt_algos_training_curves.png/.svg. -
Example Usage:
python plotting/plot_mt_algos.py
-
Script:
plotting/plot_student_policy_curves.py -
Description: Plots single-task imitation-learning histories from the per-task TensorBoard logs in
data/final_benchmarks/arnold_single_task/. -
Output:
data/figures/student_policies/(.png). -
Example Usage:
python plotting/plot_student_policy_curves.py
-
Script:
plotting/plot_transfer_vs_scratch.py -
Description: Compares learning from a pretrained multi-task policy (Transfer) against training from scratch for four downstream tasks (pen, reorient, hand_middle_reach, hand_little_reach).
-
Output:
data/figures/transfer_learning/transfer_vs_scratch_comparison.png/.svg. -
Example Usage:
python plotting/plot_transfer_vs_scratch.py
python plotting/plot_learning_curves.py --panel finetuning --raw \
--smoothing savgol --window 101 --combine_tasksReads the supplied fine-tuning cache at data/analysis/learning_curves.csv.gz and
writes figures to data/figures/learning_curves/. Other panels require exported CSVs.
See Learning curves from CSV inputs
for the export command, column definitions and multi-seed behavior.
This project is licensed under the BSD 3-Clause License.
It also includes code derived from Stable-Baselines3, imitation, PyTorch-RL, Lattice, Kinesis, PHC and IsaacGymEnvs, which remains under its own licenses (MIT, BSD 3-Clause and BSD 3-Clause Clear).
The musculoskeletal models loaded from data/kinesis/xml/ are not included in
this repository and must be obtained separately from myo_sim and Kinesis under
the Apache License 2.0.
See the LICENSE file for the full terms and per-file attributions.