简体中文 | English
This directory provides LIBERO simulation eval and latency scripts for two policy paths:
- robot.cpp C++ Policy: loads GGUF through
model-serverand runs rollout and latency checks. - LeRobot Policy: uses the original Python policy as the baseline.
eval/libero/
├── policy/
│ └── model_server.py # LIBERO observation -> C++ policy request adapter
├── runners/
│ ├── run_model_server.py # C++ policy LIBERO rollout runner
│ ├── run_lerobot.py # LeRobot Policy rollout wrapper
│ └── latency_lerobot.py # LeRobot policy latency runner
├── scripts/
│ └── run_model_server.sh # launch/eval convenience wrapper
├── utils/
│ ├── common.py # result paths, JSON writing, episode summaries
│ └── environment.py # LIBERO config, runtime env, reset/success helpers
└── environment.yaml # optional conda environment
Shared action chunk buffering lives in robot_client/policy/base_policy.py.
LIBERO C++ Policy launch, shutdown, and timing helpers live in
eval/libero/policy/model_server.py.
Using a dedicated conda environment is recommended:
conda env create -f eval/libero/environment.yaml
conda activate robotcpp-liberoIf reusing an existing environment, install at least:
pip install "cmake<4"
pip install --no-build-isolation "hf-libero>=0.1.3,<0.2.0"
pip install "lerobot[libero]"The LeRobot baseline (step4) also needs the policy's own extra. The two are
mutually exclusive (different transformers), so install per baseline:
pip install "lerobot[pi]" # pi0 (transformers fork)
pip install "lerobot[smolvla]" # smolvla (transformers>=4.57)Rendering: prefer GPU (EGL). If MUJOCO_GL=egl cannot find a device, see
Troubleshooting. OSMesa (MUJOCO_GL=osmesa,
PYOPENGL_PLATFORM=osmesa) works but is ~8x slower software rendering.
For Pi0, refer to the GGUF checkpoint
robotcpp/pi-libero-bf16.
The C++ Policy uses converted split GGUF files:
hf download robotcpp/pi-libero-bf16 \
--include "*.gguf" \
--local-dir ckpts/pi-libero-bf16The wrapper defaults to:
GGUF_DIR=ckpts/pi-libero-bf16
MODEL=pi-libero-bf16For SmolVLA, download robotcpp/smolvla-libero-bf16 and run with
MODEL_TYPE=smolvla (the wrapper then defaults GGUF_DIR=ckpts/smolvla-libero-bf16).
These repos are private/gated — use an authorized token, and set
HF_ENDPOINT=https://hf-mirror.com on restricted networks.
The simplest path is run_model_server.sh. It checks the GGUF files and an
existing model-server binary, then launches the C++ Policy and runs
eval.libero.runners.run_model_server:
CONDA_ENV=robotcpp-libero \
bash eval/libero/scripts/run_model_server.shCommon variables:
| Variable | Description |
|---|---|
MODEL_TYPE |
pi0 (default) or smolvla. |
GGUF_DIR, MODEL |
Split GGUF inputs. Defaults to the path and filename prefix from step1. |
SMOLVLA_DTYPE |
SmolVLA GGUF precision (bf16/f32; state projector stays f32). |
N_ACTION_STEPS |
Actions consumed per predicted chunk before re-querying (open-loop horizon). Default: full chunk (= chunk_size, i.e. 50 for pi0); SmolVLA defaults to 1 (closed-loop). See Action chunk execution. |
BACKEND |
C++ Policy server preset, matching robot_server/test/test_server_latency.sh. Options are linux-cuda, linux-cpu, mac-metal, and mac-cpu; default is linux-cuda. |
SERVER_BIN |
Custom model-server path. By default it is derived from BACKEND. |
HOST, PORT |
Shared client/server endpoint and must stay in sync. |
SUITE, TASK_IDS, N_EPISODES, SEED, EPISODE_LENGTH |
LIBERO rollout configuration. |
MUJOCO_GL, PYOPENGL_PLATFORM, OUTPUT |
Rendering backend and result output. |
If the default C++ Policy server does not exist, build it from the project root
README or set SERVER_BIN=/path/to/model-server directly. BUILD_DIR is only
the intermediate path used to derive the default SERVER_BIN, and usually does
not need to be set manually.
Arguments after run_model_server.sh are passed to
python -m eval.libero.runners.run_model_server, before the generated
--server-command block. For example:
OUTPUT=eval/results/pi0-libero-object.json \
bash eval/libero/scripts/run_model_server.sh --episode-length 400The output JSON contains per-episode server_timing_avg_ms and a top-level
timing_ms summary, including roundtrip_ms, server_predict_ms,
model_total_ms, prefix_ms, denoise_ms, and other metrics returned by the
C++ Policy.
If you need to write the full server command by hand, use
python -m eval.libero.runners.run_model_server --help; note that
--server-command ... must be the final part of the command.
python -m eval.libero.runners.latency_lerobot \
--policy-path lerobot/pi0_libero_finetuned_v044C++ Policy independent latency testing is unified under
robot_server/test/test_server_latency.sh.
LeRobot baseline is used to compare rollout behavior against the original Python policy:
python -m eval.libero.runners.run_lerobot \
--policy-path lerobot/pi0_libero_finetuned_v044 \
--mujoco-gl osmesa \
--pyopengl-platform osmesa \
--extra-arg=--policy.compile_model=falseThe runner writes stdout.log, stderr.log, baseline_run.json, and
LeRobot's eval_info.json under eval/results/lerobot-baseline-*.
For SmolVLA, use lerobot/smolvla_libero
(its inputs match the GGUF).
LiberoModelServerPolicy sends LIBERO observations to the C++ Policy server
with these conventions:
- LIBERO cameras
imageandimage2are sent asobservation.images.imageandobservation.images.image2. - Images are flipped along height and width to match LeRobot's
LiberoProcessorStep. - The raw 8D LIBERO state is sent to
model-serverdirectly. - After the server returns an action chunk, the policy queues it and consumes one action per environment step.
- The action sent to the LIBERO environment uses the first 7 dimensions by default.
The server returns a full action chunk (chunk_size, e.g. 50) per predict; the
client's action queue decides how many to execute before re-querying. robot.cpp
consumes the whole chunk, i.e. effectively n_action_steps = chunk_size.
This matches pi0 (native n_action_steps=50), but not SmolVLA, whose native
n_action_steps=1 (closed-loop, re-predict every step) — running SmolVLA
open-loop at 50 degrades it badly. Use N_ACTION_STEPS=1 for SmolVLA (the
wrapper's default) and match it on the baseline (--policy.n_action_steps=1).
These are common setup issues when running the eval on a fresh headless machine.
The libero package ships an empty assets/ directory, so its
get_assets_path() returns that empty path and never triggers a download. Fetch
the mesh/texture assets once into the package directory:
ASSETS=$(python -c "import libero, os; print(os.path.join(os.path.dirname(libero.__file__), 'libero', 'assets'))")
python -c "from huggingface_hub import snapshot_download; snapshot_download('lerobot/libero-assets', repo_type='dataset', local_dir='${ASSETS}')"bddl_files and init_files are bundled with the package, so only the assets
above need downloading. ensure_libero_config writes ~/.libero/config.yaml
automatically on first run.
On a headless host without /dev/dri, MUJOCO_GL=egl may fail with
Cannot initialize a EGL device display ... PLATFORM_DEVICE extension. This
usually means glvnd fell back to Mesa EGL. If the machine has an NVIDIA GPU,
register the NVIDIA EGL vendor so EGL can enumerate the GPU as a device:
sudo tee /usr/share/glvnd/egl_vendor.d/10_nvidia.json >/dev/null <<'EOF'
{ "file_format_version": "1.0.0", "ICD": { "library_path": "libEGL_nvidia.so.0" } }
EOF
# then:
export MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 # match CUDA_VISIBLE_DEVICESGPU rendering is roughly an order of magnitude faster than OSMesa software
rendering (~15 ms/step vs ~120 ms/step in our tests), which matters a lot for
full-suite rollouts. Only fall back to MUJOCO_GL=osmesa
(apt-get install -y libosmesa6, PYOPENGL_PLATFORM=osmesa) when no GPU
rendering is available.