Skip to content

Repository files navigation

Teach and Grow logo

Teach and Grow

An Agent-Centered Architecture for General Robot Learning

Chang Nie · Zhe Liu · Hesheng Wang
Shanghai Jiao Tong University

Training-free task acquisition · AI-agent reasoning · Reusable robot skills

Project website · 中文介绍 · Paper · Demonstrations

Conceptual cover: placing a plush toy into a bowl in a changed scene through sparse teaching and agent guidance

Why Teach and Grow?

A robot encounters an unfamiliar container, a changed camera view, or a grasp that no longer works. The missing capability may be local, yet adapting an end-to-end vision-language-action (VLA) or world-action model (WAM) often reopens a larger cycle of interaction collection, policy optimization, and regression testing. Physical data is costly to acquire, combinations of scene conditions are difficult to cover, and a rare failure may contain a useful lesson that should be retained explicitly.

Teach-and-Grow Learning (TGL) asks whether a robot can acquire new executable behavior while keeping its pretrained model weights fixed. A few successful demonstrations supply a starting task structure. An AI agent organizes that structure into reusable skills, grounds execution in the current scene, checks physical outcomes, and retains useful behavior and experience for later tasks.

LLMs provide the reasoning foundation. Tool-using agents such as Codex and Claude Code motivate the interaction pattern: inspect state, choose tools, act, read the result, and revise the plan. TGL brings that pattern to manipulation. The paper's concrete instantiation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex as the agent interface, with robot-native perception, planning, and control providing physical execution. The model and transport selected by a particular code run must be configured explicitly; older source defaults are not the paper's experiment configuration.

How the method works

  1. Teach a task structure. Successful demonstrations reveal shared subgoals, object roles, ordering, and observable state changes. Scene-specific poses and trajectories remain runtime variables.
  2. Build reusable Skill Blocks. Each block describes an intended effect, applicability scope, strategy, grounding function, compatible executors, success test, and recovery choices. Candidates are validated before admission to the library.
  3. Act through tools and physical feedback. The agent selects and composes skills. Executors carry out local motion, then return evidence. A failed or inconclusive effect can trigger reobservation, a different tool, or a revised plan.
  4. Grow explicit capability. The Skill Library retains executable behavior. Experience Memory retains task conditions, outcomes, diagnoses, and repairs. New tasks update these explicit stores without task-specific gradient updates, fine-tuning, or reinforcement learning.

For example, “put the bowl on the plate” preserves the acquire–transport–release structure while allowing a different grasp and motion in the current scene. The system checks that the bowl is held before continuing. The paired videos illustrate different grasps and feedback-driven recovery in LIBERO simulation.

A learned policy can still execute a Skill Block. The paper also discusses a future slow-teacher/fast-student path and a scaling hypothesis for effective reusable experience. Detailed results, protocols, and the distinction between implemented studies and proposed extensions are in the paper.

What is in this repository?

This is the public TGL repository. The current checkout contains the earlier CodeRobo / LMVS agent-and-tool implementation, the LIBERO environment, memory components, and evaluation utilities. Those names remain in Python modules for compatibility.

The public snapshot does not yet contain the complete paper-level Skill Block induction and agent-led skill-trajectory runtime. In particular, the paper's skill_card.py, atomic_skill_inducer.py, and skill_trajectory_supervisor.py modules are absent from this checkout. The entry points below exercise the published agent/tool stack; they are not a one-command reproduction of all paper results. The older notes under docs/ describe the earlier preview and should be read in that context.

Directory / entry point Published functionality
libero/ Benchmark environments, assets, BDDL tasks, initial states, and inherited learning baselines
libero_sdk/ Task discovery, environment and robot wrappers, perception and skill interfaces
lmvs/agents/codex_brain.py Structured agent decisions from task instructions and current evidence
lmvs/action_api.py Robot action interfaces, local execution, and recovery primitives
lmvs/structured_memory.py Typed experience records
lmvs/skill_registry.py Earlier skill registry, distinct from the paper's full Skill Block lifecycle
scripts/run_libero_object_sweep.py Configurable suite/task/state runner for the public agent stack
scripts/run_codex_40_task_experiment.py Experiment command orchestration, demonstration-memory extraction, and checks
scripts/ Static checks, evidence audits, focused tests, and reporting tools

Downloaded datasets, model checkpoints, local run outputs, and private model-call transcripts are not bundled with the repository.

Getting started

1. Set up the source environment

Use a dedicated Linux environment with Python, a suitable PyTorch installation, and MuJoCo rendering support. The inherited LIBERO setup originally targeted Python 3.8.13; pyproject.toml declares Python ≥3.8. The dependency manifests are not a single pinned environment lock: requirements.txt specifies Gym, while pyproject.toml additionally declares Gymnasium and MuJoCo. Check the resolved environment for the simulator and perception backends you intend to use.

git clone https://github.com/IRMVLab/TGL.git
cd TGL

conda create -n tgl python=3.8.13
conda activate tgl
python -m pip install -r requirements.txt
python -m pip install -e .

Run the following commands from the repository root. The inherited distribution name is libero; the agent modules remain under lmvs/. Optional perception services and checkpoints require their own setup.

For GPU-backed rendering, select the appropriate device for your machine:

export CUDA_VISIBLE_DEVICES=0
export MUJOCO_EGL_DEVICE_ID=0

2. Check the source and inspect an experiment plan

These checks do not run a robot rollout:

python scripts/run_lmvs_static_checks.py
python scripts/check_no_runtime_gt.py lmvs

python scripts/run_codex_40_task_experiment.py --dry-run \
  --suites libero_spatial --tasks 0 --init-states 0 \
  --codex-transport cli \
  --out-dir ./lmvs_runs \
  --demo-bundle-dir ./lmvs_demo_bundles \
  --demo-memory-dir ./lmvs_memory/demo

The dry run prints the planned extraction, execution, and audit commands. It does not validate simulator setup, model access, or task performance. Explicit output directories keep paths local to this clone.

3. Configure an agent transport

The public runner supports cli, api, subagent, and subagent_file selections, along with SDK selections that its help describes as planned stubs. Select a transport explicitly: the default subagent_file expects a separate process to answer file-based requests and is not a standalone model connection.

  • CLI: install and authenticate Codex separately, then pass --codex-transport cli. The runner invokes the configured codex executable.
  • API: provide the endpoint, credential environment variable, and a model supported by that endpoint. Pass --codex-transport api --codex-api-model YOUR_MODEL_ID. The source accepts CODEX_API_BASE_URL, CODEX_API_KEY, and CODEX_API_MODEL; its adapter uses a Chat Completions-compatible endpoint with image input.
  • File exchange: use the request/response writer workflow in scripts/run_codex_subagent_file_writer.py when an external orchestrator supplies decisions.

Credentials belong in your environment, never in tracked files. See lmvs/codex_api.py and lmvs/codex_transports/ for the implemented interfaces.

4. Run one simulation case

After checking simulator dependencies, rendering, perception backends, and model access, a small public-stack run is:

python scripts/run_libero_object_sweep.py \
  --suite libero_spatial \
  --tasks 0 --init-states 0 \
  --max-turns 6 \
  --policy-mode codex-brain \
  --codex-tool-profile generic \
  --codex-transport cli \
  --out lmvs_runs/spatial_t0_i0.json

codex-brain selects the agent policy; generic exposes generic tools rather than task-family specialists. This is a configuration example for the code present here, not a benchmark result. Inspect the output summary and associated execution evidence before interpreting success or failure.

5. Add demonstrations when needed

Demonstration datasets are separate downloads:

python benchmark_scripts/download_libero_datasets.py \
  --datasets libero_spatial --use-huggingface

The multi-task orchestration script enables demonstration-memory extraction by default. Inspect its --dry-run output before launching an experiment, and provide the required data and agent transport first. Raw observations, memory records, run directories, and generated evidence should remain outside Git.

Results and reproducibility

The project website emphasizes the research idea and qualitative demonstrations. The paper contains the numerical results and experimental protocols. Reproducing those results requires the matching implementation, model configuration, evaluation states, and run artifacts; the current public snapshot alone does not supply that complete bundle.

The code includes checks for runtime ground-truth access, replay, transcript completeness, transport health, and metadata. These checks help characterize a run; passing source checks is not evidence of robot task performance.

Citation

@techreport{nie2026teachandgrow,
  title = {Teach and Grow: An Agent-Centered Architecture for General Robot Learning},
  author = {Nie, Chang and Liu, Zhe and Wang, Hesheng},
  institution = {Shanghai Jiao Tong University},
  year = {2026},
  url = {https://tgl.changnie.top/}
}

Contact and license

Chang Nie · personal homepage · Email · Issues

See LICENSE for the repository license. Downloaded datasets, checkpoints, and third-party components may have separate terms. The logo and cover are AI-generated conceptual artwork for Teach and Grow; the cover does not depict a physical-robot experiment.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages