Chang Nie · Zhe Liu · Hesheng Wang
Shanghai Jiao Tong University
Training-free task acquisition · AI-agent reasoning · Reusable robot skills
Project website · 中文介绍 · Paper · Demonstrations
A robot encounters an unfamiliar container, a changed camera view, or a grasp that no longer works. The missing capability may be local, yet adapting an end-to-end vision-language-action (VLA) or world-action model (WAM) often reopens a larger cycle of interaction collection, policy optimization, and regression testing. Physical data is costly to acquire, combinations of scene conditions are difficult to cover, and a rare failure may contain a useful lesson that should be retained explicitly.
Teach-and-Grow Learning (TGL) asks whether a robot can acquire new executable behavior while keeping its pretrained model weights fixed. A few successful demonstrations supply a starting task structure. An AI agent organizes that structure into reusable skills, grounds execution in the current scene, checks physical outcomes, and retains useful behavior and experience for later tasks.
LLMs provide the reasoning foundation. Tool-using agents such as Codex and Claude Code motivate the interaction pattern: inspect state, choose tools, act, read the result, and revise the plan. TGL brings that pattern to manipulation. The paper's concrete instantiation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex as the agent interface, with robot-native perception, planning, and control providing physical execution. The model and transport selected by a particular code run must be configured explicitly; older source defaults are not the paper's experiment configuration.
- Teach a task structure. Successful demonstrations reveal shared subgoals, object roles, ordering, and observable state changes. Scene-specific poses and trajectories remain runtime variables.
- Build reusable Skill Blocks. Each block describes an intended effect, applicability scope, strategy, grounding function, compatible executors, success test, and recovery choices. Candidates are validated before admission to the library.
- Act through tools and physical feedback. The agent selects and composes skills. Executors carry out local motion, then return evidence. A failed or inconclusive effect can trigger reobservation, a different tool, or a revised plan.
- Grow explicit capability. The Skill Library retains executable behavior. Experience Memory retains task conditions, outcomes, diagnoses, and repairs. New tasks update these explicit stores without task-specific gradient updates, fine-tuning, or reinforcement learning.
For example, “put the bowl on the plate” preserves the acquire–transport–release structure while allowing a different grasp and motion in the current scene. The system checks that the bowl is held before continuing. The paired videos illustrate different grasps and feedback-driven recovery in LIBERO simulation.
A learned policy can still execute a Skill Block. The paper also discusses a future slow-teacher/fast-student path and a scaling hypothesis for effective reusable experience. Detailed results, protocols, and the distinction between implemented studies and proposed extensions are in the paper.
This is the public TGL repository. The current checkout contains the earlier CodeRobo / LMVS agent-and-tool implementation, the LIBERO environment, memory components, and evaluation utilities. Those names remain in Python modules for compatibility.
The public snapshot does not yet contain the complete paper-level Skill Block induction and agent-led skill-trajectory runtime. In particular, the paper's skill_card.py, atomic_skill_inducer.py, and skill_trajectory_supervisor.py modules are absent from this checkout. The entry points below exercise the published agent/tool stack; they are not a one-command reproduction of all paper results. The older notes under docs/ describe the earlier preview and should be read in that context.
| Directory / entry point | Published functionality |
|---|---|
libero/ |
Benchmark environments, assets, BDDL tasks, initial states, and inherited learning baselines |
libero_sdk/ |
Task discovery, environment and robot wrappers, perception and skill interfaces |
lmvs/agents/codex_brain.py |
Structured agent decisions from task instructions and current evidence |
lmvs/action_api.py |
Robot action interfaces, local execution, and recovery primitives |
lmvs/structured_memory.py |
Typed experience records |
lmvs/skill_registry.py |
Earlier skill registry, distinct from the paper's full Skill Block lifecycle |
scripts/run_libero_object_sweep.py |
Configurable suite/task/state runner for the public agent stack |
scripts/run_codex_40_task_experiment.py |
Experiment command orchestration, demonstration-memory extraction, and checks |
scripts/ |
Static checks, evidence audits, focused tests, and reporting tools |
Downloaded datasets, model checkpoints, local run outputs, and private model-call transcripts are not bundled with the repository.
Use a dedicated Linux environment with Python, a suitable PyTorch installation, and MuJoCo rendering support. The inherited LIBERO setup originally targeted Python 3.8.13; pyproject.toml declares Python ≥3.8. The dependency manifests are not a single pinned environment lock: requirements.txt specifies Gym, while pyproject.toml additionally declares Gymnasium and MuJoCo. Check the resolved environment for the simulator and perception backends you intend to use.
git clone https://github.com/IRMVLab/TGL.git
cd TGL
conda create -n tgl python=3.8.13
conda activate tgl
python -m pip install -r requirements.txt
python -m pip install -e .Run the following commands from the repository root. The inherited distribution name is libero; the agent modules remain under lmvs/. Optional perception services and checkpoints require their own setup.
For GPU-backed rendering, select the appropriate device for your machine:
export CUDA_VISIBLE_DEVICES=0
export MUJOCO_EGL_DEVICE_ID=0These checks do not run a robot rollout:
python scripts/run_lmvs_static_checks.py
python scripts/check_no_runtime_gt.py lmvs
python scripts/run_codex_40_task_experiment.py --dry-run \
--suites libero_spatial --tasks 0 --init-states 0 \
--codex-transport cli \
--out-dir ./lmvs_runs \
--demo-bundle-dir ./lmvs_demo_bundles \
--demo-memory-dir ./lmvs_memory/demoThe dry run prints the planned extraction, execution, and audit commands. It does not validate simulator setup, model access, or task performance. Explicit output directories keep paths local to this clone.
The public runner supports cli, api, subagent, and subagent_file selections, along with SDK selections that its help describes as planned stubs. Select a transport explicitly: the default subagent_file expects a separate process to answer file-based requests and is not a standalone model connection.
- CLI: install and authenticate Codex separately, then pass
--codex-transport cli. The runner invokes the configuredcodexexecutable. - API: provide the endpoint, credential environment variable, and a model supported by that endpoint. Pass
--codex-transport api --codex-api-model YOUR_MODEL_ID. The source acceptsCODEX_API_BASE_URL,CODEX_API_KEY, andCODEX_API_MODEL; its adapter uses a Chat Completions-compatible endpoint with image input. - File exchange: use the request/response writer workflow in
scripts/run_codex_subagent_file_writer.pywhen an external orchestrator supplies decisions.
Credentials belong in your environment, never in tracked files. See lmvs/codex_api.py and lmvs/codex_transports/ for the implemented interfaces.
After checking simulator dependencies, rendering, perception backends, and model access, a small public-stack run is:
python scripts/run_libero_object_sweep.py \
--suite libero_spatial \
--tasks 0 --init-states 0 \
--max-turns 6 \
--policy-mode codex-brain \
--codex-tool-profile generic \
--codex-transport cli \
--out lmvs_runs/spatial_t0_i0.jsoncodex-brain selects the agent policy; generic exposes generic tools rather than task-family specialists. This is a configuration example for the code present here, not a benchmark result. Inspect the output summary and associated execution evidence before interpreting success or failure.
Demonstration datasets are separate downloads:
python benchmark_scripts/download_libero_datasets.py \
--datasets libero_spatial --use-huggingfaceThe multi-task orchestration script enables demonstration-memory extraction by default. Inspect its --dry-run output before launching an experiment, and provide the required data and agent transport first. Raw observations, memory records, run directories, and generated evidence should remain outside Git.
The project website emphasizes the research idea and qualitative demonstrations. The paper contains the numerical results and experimental protocols. Reproducing those results requires the matching implementation, model configuration, evaluation states, and run artifacts; the current public snapshot alone does not supply that complete bundle.
The code includes checks for runtime ground-truth access, replay, transcript completeness, transport health, and metadata. These checks help characterize a run; passing source checks is not evidence of robot task performance.
@techreport{nie2026teachandgrow,
title = {Teach and Grow: An Agent-Centered Architecture for General Robot Learning},
author = {Nie, Chang and Liu, Zhe and Wang, Hesheng},
institution = {Shanghai Jiao Tong University},
year = {2026},
url = {https://tgl.changnie.top/}
}Chang Nie · personal homepage · Email · Issues
See LICENSE for the repository license. Downloaded datasets, checkpoints, and third-party components may have separate terms. The logo and cover are AI-generated conceptual artwork for Teach and Grow; the cover does not depict a physical-robot experiment.