Homepage

About Me

I am currently a master's student in Computer Application Technology at the University of Chinese Academy of Sciences. My work focuses on large language model training and evaluation, including SFT, RLHF/RLVR, reward modeling, rubric-based reinforcement learning, and agentic evaluation.

I have worked on AI Office agentic evaluation, safety expert model training, customer-service agent training, and multimodal/audio-language model research. I am also interested in practical training insights around KL placement, rollout staleness, off-policy effects, multi-reward optimization, online policy distillation, and token-level credit assignment.

LLM Training RLHF / RLVR Agentic Evaluation Rubric-based RL Multi-Reward RL Online Policy Distillation Token-Level Credit Audio-Language Model Multimodal LLM

Updates

News

  • SMOPD and EMPIRE were released on arXiv and submitted to AAAI 2027.
  • Open Rubric System was accepted by EMNLP 2026.
  • CoRT was released on arXiv and submitted to AAAI 2027.
  • Towards Region-Level No-Reference Image Quality Assessment was accepted by IEEE TIP.
  • Joined Alibaba Qwen as an algorithm engineer intern.
  • HumanPCR was accepted as an ICLR 2026 poster.
  • Joined ByteDance Data e-commerce customer-service NLP team as an algorithm engineer intern.
  • MATS was accepted by ICML 2025.
  • Received B.E. in Software Engineering from Beijing Jiaotong University.

Research

Publications

Published / Accepted

MATS architecture ICML 2025

MATS: An Audio Language Model under Text-only Supervision

Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen

International Conference on Machine Learning, 2025 · First author

An audio language model trained under text-only supervision, with Santa for transferring audio embeddings into the text embedding space while preserving audio semantics.

EMNLP 2026

Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric

Wen Wang, Ruipeng Jia, Yunyi Yang, Yuxin Wu, Yongbo Gai, Siyuan Tao, Mengyu Zhou, Jianhe Lin, Xiaoxi Jiang, Guanjun Jiang

Empirical Methods in Natural Language Processing, 2026 · First author

A rubric-based LLM-as-a-Judge reward system with Pairwise Adaptive Meta-Rubrics and Pointwise Verifiable Rubrics for open-domain reward supervision.

Region-level no-reference image quality assessment framework IEEE TIP

Towards Region-Level No-Reference Image Quality Assessment

Zewen Chen, Juan Wang, Wen Wang, Sunhan Xu, Hang Xiong, Yun Zeng, Jian Guo, Shuxun Wang, Chunfeng Yuan, Bing Li, Weiming Hu

IEEE Transactions on Image Processing, 2026 · Accepted

Region-level no-reference image quality assessment with component-level and object-level mask data, local quality scores, and a mask-based feature extractor.

HumanPCR taxonomy, annotation, and quality control framework ICLR 2026

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

Keliang Li, Hongze Shen, Hao Shi, RuiBing Hou, Hong Chang, Jie Huang, Chenghao Jia, Wen Wang, Yiling Wu, Dongmei Jiang, Shiguang Shan, Xilin Chen

International Conference on Learning Representations, 2026 · Poster

A benchmark for probing multimodal large language models across human-centric perception, comprehension, and reasoning tasks.

Preprints / Under Review

Work

Experience

2026.04 - Present

Alibaba | Qwen Business Unit, LLM Training and Application Algorithm Foundation Group · Algorithm Engineer Intern

  • Worked on AI Office agentic evaluation and safety expert model training, covering benchmark construction, data flywheel, SFT data generation, model training, and automated evaluation.
  • Built GDPVal-style long-horizon office evaluation and trajectory analysis, participated in Qwen3.5-Plus/Qwen3.5-397B-A17B safety expert model training, and improved checklist-human agreement from 67.38% to 99.47%.
2025.10 - 2026.03

ByteDance | Data E-commerce Platform Customer Service NLP LLM Team · Algorithm Engineer Intern

Worked on full-scenario customer-service ChatBot agent training and evaluation, including SFT, RLHF, data cleaning, distillation data construction, model training, and online effect tracking.

Trained Qwen3-32B with ms-swift, FSDP, and Megatron-LM, and built multi-stage RLHF workflows with point-wise and pair-wise reward models.

Background

Education

2025.09 - 2028.06

University of Chinese Academy of Sciences

M.S. student, Computer Application Technology.

2021.09 - 2025.06

Beijing Jiaotong University

B.E. in Software Engineering. Ranked 3/169, received three special scholarships, and was named an Excellent Graduate of Beijing and Beijing Jiaotong University.

Honors

Honors and Awards

  • First Prize, National College Student Mathematics Competition, non-mathematics group.
  • CSP 340, cumulative top 1.98%.
  • Excellent Graduate of Beijing.
  • First Prize, Beijing College Student Mathematics Competition.
  • Second Prize, 2023 Lanqiao Cup Provincial Contest, C++ A group.
  • Second Prize, Regional Contest of National College Student Software Innovation Competition.