date: 2026.8 标题: VLA、WAM 之外:OpenETA 如何把智能体闭环带入物理世界 title: OpenETA: Bringing the Agentic Loop into the Physical World 描述: OpenETA 将数字智能体的“理解任务—调用工具—接收反馈—继续决策”闭环带入仿真器和机器人,以可信回执、新观测义务和可回放轨迹连接感知、行动与验证。 desc: OpenETA brings the digital-agent loop of understanding, tool use, feedback, and continued decision-making into simulators and robots through trusted receipts, fresh-observation obligations, and replayable trajectories. 标签: 具身智能, 智能体, 机器人 tags: Embodied AI, Agents, Robotics 图: assets/img/highlights/openeta.jpg 链接: /blog/cn/openeta/
date: 2026.7 标题: MOSS-VL:面向实时视频流的开源视觉语言模型 title: MOSS-VL: Open-Weight Vision-Language Models for Real-Time Video Streams 描述: OpenMOSS 发布的视觉语言模型系列,覆盖持续视频流实时交互、离线长视频理解与领域微调;本文进一步对比 GPT-4o、Gemini、Qwen2.5-VL、InternVL3 和 LLaVA-OneVision 的选型边界。 desc: OpenMOSS's vision-language model family for continuous video-stream interaction, offline long-video understanding, and domain fine-tuning, with a practical positioning guide against GPT-4o, Gemini, Qwen2.5-VL, InternVL3, and LLaVA-OneVision. 标签: 多模态, 视频, 开源模型 tags: Multimodal, Video, Open Models 图: assets/img/highlights/moss-vl-hero.png 链接: /blog/cn/moss-vl/
date: 2026.6 标题: 组织智能:在组织层面治理 Agentic AI title: Organizational Intelligence: Governing Agentic AI at the Level of the Organization 描述: 当 AI 从「模型」走向「自主智能体」,真正的瓶颈从模型转移到它所处的组织。本文主张:组织——而非孤立的模型或任务——才是设计、评估与治理 Agentic AI 的正确单位,并给出八要素组织状态形式化、三层嵌套反馈循环,以及以治理为准入条件的 L0–L5 成熟度模型。 desc: As AI moves from capable models to autonomous agents, the binding constraint shifts from the model to the organization around it. This essay argues the organization—not the isolated model or task—is the right unit for designing, evaluating, and governing agentic AI, with an eight-element state formalization, three nested feedback loops, and a governance-as-entry-condition L0–L5 maturity model. 标签: AI4AI, 智能体 tags: AI4AI, Agents 图: assets/img/highlights/organizational-intelligence.webp 链接: /blog/en/organizational-intelligence/
date: 2026.6 标题: Thinking with Video:用视频生成做多模态推理 title: Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 描述: 提出“用视频思考”范式——让 Sora-2 等视频生成模型以视频帧为统一媒介进行多模态推理,弥补文字/图像难以刻画动态过程的不足;配套 VideoThinkBench 上视觉与文本推理均表现强劲,视觉谜题超越 GPT-5。 desc: A “Thinking with Video” paradigm—video-generation models like Sora-2 reason over generated video frames as a unified medium, capturing the dynamic processes that text/image “thinking” cannot. On the new VideoThinkBench it rivals top VLMs and surpasses GPT-5 on visual puzzles. 标签: 多模态, 视频 tags: Multimodal, Video 图: assets/img/highlights/thinking-with-video.webp 链接: /blog/cn/thinking-with-video/
date: 2026.3 标题: MOSS-TTS 技术报告 title: MOSS-TTS Technical Report 描述: 可扩展的语音生成基座模型,支持零样本音色克隆、时长与发音控制、流畅中英混说与长语音生成。 desc: A scalable speech-generation foundation model supporting zero-shot voice cloning, duration and pronunciation control, smooth code-switching, and stable long-form generation. 标签: 语音 tags: Speech 图: assets/img/highlights/moss-tts.webp 链接: /blog/cn/moss-tts/
date: 2026.3 标题: AI 也能学会“科学品味” title: AI Can Learn Scientific Taste 描述: 提出“社区反馈强化学习”(RLCF):用大规模引用信号训练“科学评审”判断创意价值,并对齐出“科学思考者”提出高影响力研究构想,判断力超越 GPT-5.2、Gemini 3 Pro。 desc: Introduces Reinforcement Learning from Community Feedback (RLCF): a "Scientific Judge" trained on large-scale citation signals to assess idea quality, and a "Scientific Thinker" aligned to propose high-impact research ideas—surpassing frontier models like GPT-5.2 and Gemini 3 Pro. 标签: AI4AI tags: AI4AI 图: assets/img/highlights/scientific-taste.webp 链接: /blog/cn/scientific-taste/
date: 2026.2 标题: MOVA:可扩展的同步视频-音频生成 title: MOVA: Towards Scalable and Synchronized Video-Audio Generation 描述: 开源视频-音频联合生成模型,可同步生成高质量画面与声音——唇形同步语音、环境音效与内容匹配的音乐;采用 MoE 架构(总参 32B,激活 18B)。 desc: An open-source joint video-audio generation model producing high-quality, synchronized visuals and audio—lip-synced speech, environment-aware sound effects, and content-aligned music—built on a MoE architecture (32B total / 18B active). 标签: 视频, 语音 tags: Video, Speech 图: assets/img/highlights/mova.webp 链接: /blog/cn/mova/
date: 2025.10 标题: MOSS-Speech: 真语音到语音生成 title: MOSS-Speech: True Speech-to-Speech Generation 描述: 原生端到端语音交互,无需任何中间文本引导 desc: Native end-to-end speech interaction without any intermediate text guidance 标签: 语音 tags: Speech 图: assets/img/highlights/moss-speech.webp 链接: https://arxiv.org/abs/2510.00499
date: 2025.6 标题: XY-Tokenizer: 低码率声学语义统一编码 title: XY-Tokenizer: Low-Bitrate Unified Acoustic-Semantic Encoding 描述: 1kbps最强声学语义统一编码及离散化工具 desc: State-of-the-art 1kbps unified acoustic-semantic encoding and discretization tool 标签: 语音 tags: Speech 图: assets/img/highlights/xy-tokenizer.webp 链接: https://arxiv.org/abs/2506.23325