Skip to content
View sylvain-wei's full-sized avatar
😎
Hahaha
😎
Hahaha

Highlights

  • Pro

Block or report sylvain-wei

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sylvain-wei/README.md

hi, here's Shaohang Wei (魏少杭) Profile views

personal page / cv · scholar · email · X · rednote

I am currently a Ph.D. student at Peking University.

I build models:

  • Post-train: Data-efficient and stable reinforcement learning across domains, including RL and OPD for agentic AI and reasoning.
  • Data & Eval: Data synthesis research and reliable evaluation.

Open for any intern opportunities.

Pinned Loading

  1. TIME TIME Public

    [NeurIPS 2025 (Spotlight✨)] TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenario

    Python 33

  2. 24-Game-Reasoning 24-Game-Reasoning Public

    超简单复现Deepseek-R1-Zero和Deepseek-R1,以「24点游戏」为例。通过zero-RL、SFT以及SFT+RL,以激发LLM的自主验证反思能力。 About Clean, minimal, accessible reproduction of DeepSeek R1-Zero, DeepSeek R1

    Python 35 2

  3. BloodArena/BloodArena BloodArena/BloodArena Public

    JavaScript 24 6

  4. ReVuE ReVuE Public

    [UnderReview'27] On-Policy Visual Evidence Distillation

    Python 2

  5. VISR VISR Public

    [UnderReview'27] Verifier-Induced Support Reshaping in On-Policy Optimization

    Python 2

  6. NightingaleCen/LeafyLingo NightingaleCen/LeafyLingo Public

    An interactive plant recognition system with multi-model knowledge graph

    Python 1