MarkDiffusion accepted by the Journal of Machine Learning Research (JMLR).
Aiwei Liu 刘瑷玮
Model Architecture & Pretraining Infrastructure
Model architecture.
Pretraining at scale.
I build large-scale foundation models at WeChat AI, Tencent. As a core member of the WeLM pretraining team, I define WeLM’s model architecture and develop pretraining infrastructure. I am a core contributor to WeLM-80B and WeLM-617B, and the first contributor to WeDLM.
I received my Ph.D. from Tsinghua University in 2025. My research has received 3,000+ Google Scholar citations. I have authored 20+ first-author or corresponding-author papers, including oral and spotlight presentations at leading machine learning conferences such as ICML and ICLR.
News.
Locally Confident, Globally Stuck accepted to COLM 2026.
Released WeLM-617B and Hidden Decoding at Scale and its accompanying WeLM technical blog.
WeDLM accepted to ICML 2026.
Three papers accepted to ICLR 2026.
Released WeDLM, with open-source code and model checkpoints.
Joined the WeLM team at WeChat AI, Tencent as a researcher.
Received my Ph.D. from Tsinghua University and the Outstanding Graduate of Beijing distinction.
Ideas into models.
WeLM foundation models
Core contributor · WeLM-80B & WeLM-617B
As a core member of the WeLM pretraining team, I am responsible for defining WeLM’s model architecture and work on large-scale pretraining infrastructure, connecting model design, training systems, and inference efficiency.
WeLM-80B
- Architecture
- 80B total parameters · 3B active parameters in a sparse mixture-of-experts model.
- Training & capabilities
- Fewer than 14T training tokens · 128K context, with strong reasoning and multilingual performance.
WeLM-617B
- Model architecture
- Define WeLM’s model structure with a focus on model quality and inference efficiency.
- Pretraining infrastructure
- Develop large-scale training systems and infrastructure for foundation model pretraining at the 617B scale.
WeDLM
Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
First contributor · First author
WeDLM brings diffusion language modeling to standard causal attention, enabling parallel generation while retaining the attention structure used by efficient inference systems.
Hidden Decoding at Scale
First author · Foundation model capability scaling
A sequence-length scaling method that expands latent computation for large language models, demonstrated at 100B+ MoE scale.
- Evaluated scale
- 80B and 617B model baselines.
- Key result
- Improvements across all nine evaluated benchmarks, extending sequence-length scaling to large MoE models.
Research by direction.
Representative work and first-author papers, grouped by research area. † Corresponding author.
Foundation models & efficient generation
LLM alignment
Token-level importance weighting for direct preference optimization.
Trustworthy LLMs & watermarking
Project lead for an open-source toolkit unifying 20 LLM watermarking methods, 12 evaluation tools, and automated evaluation pipelines.
Semantic-invariant watermarking for robust provenance of LLM-generated text.
Publicly verifiable watermarking for language models.
An open-source toolkit for generative watermarking of latent diffusion models.
Earlier work: text-to-SQL & NLP robustness
Research background.
EXPERIENCE
JUL 2025 — PRESENT
WeChat AI, Tencent
Researcher · WeLM Team
Core member of the WeLM pretraining team. Responsible for model architecture design; specializing in pretraining infrastructure.
Apple AIML
Research Intern · Advisor: Dr. Meng Cao
CUHK MISC Lab
Visiting Scholar · Advisor: Prof. Irwin King
UIC BDSC Lab
Visiting Scholar · Advisor: Prof. Philip S. Yu
EDUCATION
2020 — 2025
Tsinghua University
Ph.D. in Software Engineering
Advisor: Prof. Lijie Wen
Research in LLM alignment and
trustworthy LLMs.
2016 — 2020
Nanjing University
B.E. in Software Engineering
Better models.
Bigger possibilities.
Open to research collaborations in foundation model architecture,
large-scale pretraining, and efficient generation.