Skip to content
@THUDM

THUKEG

ChatGLM, GLM-4, CogVLM, CodeGeeX, CogView, ImageReward, CogVideoX | CogDL, GraphMAE, AMiner | Zhipu.ai (Z.ai) & Knowledge Engineering Group (KEG)

Pinned Loading

  1. GLM GLM Public

    GLM (General Language Model)

    Python 3.7k 373

  2. slime slime Public

    slime is an LLM post-training framework for RL Scaling.

    Python 8.3k 1.2k

  3. P-tuning-v2 P-tuning-v2 Public

    An optimized deep prompt tuning strategy comparable to fine-tuning across scales and tasks

    Python 2.1k 212

  4. ReST-MCTS ReST-MCTS Public

    ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search (NeurIPS 2024)

    Python 711 51

  5. T1 T1 Public

    RL Scaling and Test-Time Scaling (ICML'25)

    116 1

  6. AgentRL AgentRL Public

    Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework

    Python 347 28

Repositories

Showing 10 of 130 repositories
  • slime Public

    slime is an LLM post-training framework for RL Scaling.

    Python 8,313 Apache-2.0 1,200 219 233 Updated Aug 28, 2026
  • SCALE-CUA Public

    Open-source framework for computer use agents: VeriGen verifiable task synthesis, online RL training (AgentRL), and OSWorld/ScienceBoard evaluation.

    THUDM/SCALE-CUA's past year of commit activity
    Python 55 2 4 0 Updated Aug 3, 2026
  • KARL Public
    THUDM/KARL's past year of commit activity
    Python 4 1 0 0 Updated Jul 7, 2026
  • CodeRM-NT Public

    [Findings of ACL 2026] CodeRM-NT: Reward Model for Code RL without Unit Tests

    THUDM/CodeRM-NT's past year of commit activity
    Python 0 0 0 0 Updated Jul 1, 2026
  • DeepDive Public

    DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL

    THUDM/DeepDive's past year of commit activity
    Python 345 38 0 0 Updated Jun 17, 2026
  • INFTY Public

    INFTY Engine: An Optimization Toolkit to Support Continual AI

    THUDM/INFTY's past year of commit activity
    Python 574 MIT 11 6 0 Updated Jun 8, 2026
  • ReST-RL Public

    Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding

    THUDM/ReST-RL's past year of commit activity
    Python 18 MIT 0 0 0 Updated May 6, 2026
  • CaRR Public

    This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards".

    THUDM/CaRR's past year of commit activity
    Python 73 MIT 9 0 0 Updated Apr 8, 2026
  • IndexCache Public

    IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

    THUDM/IndexCache's past year of commit activity
    137 11 7 1 Updated Mar 14, 2026
  • AgentBench Public

    A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

    THUDM/AgentBench's past year of commit activity
    Python 3,703 Apache-2.0 278 64 (40 issues need help) 12 Updated Feb 8, 2026