Skip to content

About

面向 ALFWorld 长程任务规划与探索的结构化记忆 LLM 智能体。

Resources

Stars

10 stars

Watchers

0 watching

Forks

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

StructMem Agent:基于结构化记忆的 ALFWorld 文本智能体

StructMem Agent(Structured Memory Agent)基于 ALFWorld 的 TextWorld 文本环境,实现了一个 observation -> structured memory -> LLM action 的原型智能体。核心目标是把自然语言 observation 转换成可累积的结构化记忆,再让 LLM 基于这份记忆选择下一步动作。

应用场景

StructMem Agent 主要面向需要连续观察、记忆环境状态并执行多步操作的文本交互任务,当前适合以下应用和研究场景:

  • 家庭任务规划:在 ALFWorld 模拟环境中完成寻找、拿取、清洁、加热、冷却和放置物品等多步家庭任务。
  • 长程交互决策:利用对象关系、访问状态和动作历史减少遗忘与重复探索,为需要多轮环境交互的任务选择下一步动作。
  • LLM Agent 记忆研究:比较原始 observation、完整历史和结构化对象图等不同上下文组织方式对决策效果的影响。
  • 探索策略实验:通过 free_llm 与 priority_exploration 模式研究自由决策、优先级引导和重复访问抑制机制。
  • Agent 评测与错误分析:记录完整 episode 轨迹、失败案例和记忆状态,用于分析结构解析错误、无效动作、重复探索和任务失败原因。
  • 具身智能原型验证:在成本较低的文本环境中验证高层任务规划与记忆机制,为后续连接视觉感知或真实机器人执行模块提供实验基础。

当前实现聚焦 ALFWorld 的 TextWorld 环境,属于研究与实验原型,并非可直接部署到真实家庭环境的机器人控制系统。

ALFWorld 背景

ALFWorld 是论文 ALFWorld: Aligning Text and Embodied Environments for Interactive Learning 对应的交互式学习环境。它将 TextWorld 文本环境与 ALFRED 具身任务对齐,使 agent 可以先在抽象文本空间中学习和推理高层策略,再迁移到具身环境中执行任务。

本仓库聚焦 ALFWorld 的文本环境 AlfredTWEnv。THOR 视觉环境通常需要额外图形环境支持,不是当前自定义 agent 的主要开发对象。

本项目实现了什么

本项目在 ALFWorld 原始环境之上实现了一套结构化记忆驱动的自定义 agent 框架。整体流程如下:

Agent framework

核心执行链路是:

ALFWorld Environment
  -> StructContributeAgentLLM
  -> Shared Memory
  -> ActionAgent
  -> Environment Step

当前版本是 LLM-only 的结构:

  • StructContributeAgentLLM 负责结构化解析 observation。
  • ActionAgent 负责根据 Memory 和 admissible actions 选择动作。
  • 动作决策模式只有 free_llm 和 priority_exploration。

核心模块

AgentFramework

AgentFramework 是总控模块,负责串联环境、结构化解析器、共享记忆和动作决策器。它在每个 episode 中执行:

reset -> parse -> decide -> step -> repeat

它还负责评测多局 episode、记录失败轨迹,并生成 JSON / Markdown 评测报告。

StructContributeAgentLLM

StructContributeAgentLLM 读取当前 observation 和 admissible actions,让 LLM 输出结构化 JSON,提取:

  • 当前任务 task
  • 当前房间 room_name
  • 可见物体 visible_objects
  • 物体类型 room / container / item
  • 父子关系 parent_name
  • 访问状态 visited

Memory / Object

Memory 是整套系统的共享状态中心。它用 Object 节点表示房间、容器和普通物品,并维护对象图、动作历史、最近 observation、最近 parser result 和优先级信息。

Memory module

Memory 的作用不是替代 LLM,而是为 LLM 提供更稳定的中间状态表示。它把环境从“一段自然语言”整理成可读写的对象图,例如:

kitchen -> fridge -> apple

同时,Memory 会根据动作目标更新访问状态、移动父子关系,并对同名同级物体进行优先级衰减,减少连续翻找大量相似容器的情况。

ActionAgent

ActionAgent 基于 Memory 构造 prompt,让 LLM 从 admissible actions 中原样选择下一步动作。模型输出格式为:

{
  "thought": "short reasoning",
  "action": "one exact action from the admissible actions"
}

当前支持两种 LLM prompt mode:

  • free_llm: 优先级排序只作为参考,主要依赖 LLM 自由决策。
  • priority_exploration: 显式引导 LLM 优先考虑高优先级动作,减少重复探索。

Action modes

priority_exploration_epsilon 控制每一步采用 priority-guided prompt 的概率;剩余概率走 free_llm。

环境与数据配置

<your-project-root> 表示本仓库根目录。

1. 准备 Python 环境

推荐使用现有 conda 环境:

conda activate StructMemAgent

如果还没有把 alfworld 安装成可编辑包,可以执行:

cd <your-project-root>/alfworld
pip install -e .

检查环境是否可用:

cd <your-project-root>/alfworld
python -c "import alfworld; from alfworld.agents.environment import get_environment; print('alfworld import ok')"

2. 下载并配置 ALFWorld 数据

推荐显式设置 ALFWORLD_DATA,把 ALFWorld 数据放到一个可复用的位置:

export ALFWORLD_DATA=<your-project-root>/alfworld_data

也可以换成任意有足够空间的目录:

export ALFWORLD_DATA=/path/to/alfworld_data

下载 ALFWorld 的 PDDL、游戏文件和相关资源:

cd <your-project-root>/alfworld
python scripts/alfworld-download

ALFWorld 上游也支持安装命令 alfworld-download。如果不设置 ALFWORLD_DATA,上游工具可能会使用用户缓存目录;为了复现实验,建议始终显式设置该变量。

可选检查:

ls "$ALFWORLD_DATA/json_2.1.1"
ls "$ALFWORLD_DATA/logic"

LLM API 配置

根目录 .env 约定如下:

OPENAI_API_KEY=
OPENAI_BASE_URL=
  • 使用官方 OpenAI API 时,通常只需要填写 OPENAI_API_KEY。
  • 使用兼容 OpenAI 协议的第三方服务时,再额外设置 OPENAI_BASE_URL。
  • 不要把真实 API key 提交到版本库。

如何运行 my_agent

先进入环境和项目目录:

conda activate StructMemAgent
export ALFWORLD_DATA=<your-project-root>/alfworld_data
cd <your-project-root>/alfworld

如果你的数据放在其他位置,就把 ALFWORLD_DATA 换成对应目录。

快速测试

python my_agent/test_agent/run_framework.py --episodes 1 --max-steps 20

完整示例

python my_agent/test_agent/run_framework.py \
  --config configs/base_config.yaml \
  --struct-model deepseek-v4-flash \
  --action-model deepseek-v4-flash \
  --train-eval eval_in_distribution \
  --episodes 10 \
  --max-steps 25 \
  --priority-epsilon 0.4

关闭逐步日志

python my_agent/test_agent/run_framework.py --episodes 1 --max-steps 20 --quiet

参数说明

参数 含义
--config ALFWorld YAML 配置文件路径,默认 configs/base_config.yaml
--episodes 评测 episode 数量
--max-steps 每个 episode 最大步数
--train-eval 数据划分,可选 train、eval_in_distribution、eval_out_of_distribution
--struct-model 结构化解析模型
--action-model 动作决策模型
--priority-epsilon 使用 priority_exploration prompt 的概率
--quiet 关闭逐步日志

当前入口固定使用 LLM 结构化解析和 LLM 动作决策。

评测输出与代表性结果

run_framework.py 在运行前会打印完整运行配置,结束后会输出:

  • 平均 reward
  • 平均步数
  • 成功率
  • 成功局数与失败局数
  • JSON 报告路径
  • Markdown 报告路径

失败 episode 会追加记录到:

alfworld/my_agent/logs/failed_episodes.jsonl

评测报告默认输出到:

alfworld/my_agent/logs/eval_reports/

一个代表性的历史运行快照如下:

报告 数据划分 episodes max steps epsilon success rate avg steps
eval_report_20260525-113138.json eval_in_distribution 10 25 0.4 100% 11.9

该结果是历史运行快照。LLM 模型版本、API 服务、采样行为和环境状态变化都可能影响复现结果。

当前限制与后续方向

  • LLM JSON 输出仍可能不稳定;当前实现会在动作无效时直接暴露错误,而不是静默兜底。
  • Memory 使用单例模式,目前适合 batch_size=1 的顺序评测,不适合直接并行多局。
  • 目前缺乏大规模的消融实验,证明Memory模块、优先级机制等方法的有效性。
  • 后续可以继续增强目标跟踪、失败案例分析、学习化优先级和更系统的评测对比。
  • 后续也许会作为个人科研项目,进行进一步深入探索。

参考文献

以下工作为本项目所使用的环境,以及结构化记忆、推理与行动和具身智能体设计提供了主要参考:

  1. Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., & Hausknecht, M. (2021). ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. International Conference on Learning Representations (ICLR). 论文 · 项目主页
  2. Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., & Fox, D. (2020). ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 论文
  3. Côté, M.-A., Kádár, Á., Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Hausknecht, M., El Asri, L., Adada, M., Tay, W., & Trischler, A. (2018). TextWorld: A Learning Environment for Text-based Games. 论文
  4. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. International Conference on Learning Representations (ICLR). 论文 · 项目主页
  5. Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. Advances in Neural Information Processing Systems (NeurIPS). 论文
  6. Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. ACM Symposium on User Interface Software and Technology (UIST). 论文
  7. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. 论文 · 项目主页

About

面向 ALFWorld 长程任务规划与探索的结构化记忆 LLM 智能体。

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages