MarkDiffusion accepted by the Journal of Machine Learning Research (JMLR).
Aiwei Liu 刘瑷玮
Model Architecture & Pretraining Infrastructure
Model architecture.
Pretraining at scale.
I build large-scale foundation models at WeChat AI, Tencent. As a core member of the WeLM pretraining team, I define WeLM’s model architecture and develop pretraining infrastructure. I am a core contributor to WeLM-80B and WeLM-600B, and the first contributor to WeDLM.
I received my Ph.D. from Tsinghua University in 2025. My research has received 3,000+ Google Scholar citations. I have authored 20+ first-author or corresponding-author papers, including oral and spotlight presentations at leading machine learning conferences such as ICML and ICLR.
RECENT MILESTONES
News.
Locally Confident, Globally Stuck accepted to COLM 2026.
Released Hidden Decoding at Scale and its accompanying WeLM technical blog.
WeDLM accepted as an ICML 2026 Oral paper.
Three papers accepted to ICLR 2026.
Released WeDLM, with open-source code and model checkpoints.
Joined the WeLM team at WeChat AI, Tencent as a researcher.
Received my Ph.D. from Tsinghua University and the Outstanding Graduate of Beijing distinction.
FOUNDATION MODELS & RESEARCH
Ideas into models.
WeLM foundation models
Core contributor · WeLM-80B & WeLM-600B
As a core member of the WeLM pretraining team, I am responsible for defining WeLM’s model architecture and work on large-scale pretraining infrastructure, connecting model design, training systems, and inference efficiency.
WeLM-80B
CORE CONTRIBUTOR
- Architecture
- 80B total parameters · 3B active parameters in a sparse mixture-of-experts model.
- Training & capabilities
- Fewer than 14T training tokens · 128K context, with strong reasoning and multilingual performance.
WeLM-600B
CORE CONTRIBUTOR
- Model architecture
- Define WeLM’s model structure with a focus on model quality and inference efficiency.
- Pretraining infrastructure
- Develop large-scale training systems and infrastructure for foundation model pretraining at the 600B scale.
WeDLM
Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
First contributor · First author
WeDLM brings diffusion language modeling to standard causal attention, enabling parallel generation while retaining the attention structure used by efficient inference systems.
Reported speedups against optimized vLLM baselines, depending on the generation workload.
LATENT COMPUTATION · 2026
Hidden Decoding at Scale
First author · Foundation model capability scaling
A sequence-length scaling method that expands latent computation for large language models, demonstrated at 100B+ MoE scale.
- Evaluated scale
- 80B and 617B model baselines.
- Key result
- Improvements across all nine evaluated benchmarks, extending sequence-length scaling to large MoE models.
SELECTED PUBLICATIONS
Research by direction.
Representative work and first-author papers, grouped by research area.
01Foundation models & efficient generation
02LLM alignment
Token-level importance weighting for direct preference optimization.
03Trustworthy LLMs & watermarking
An open-source toolkit unifying 20 LLM watermarking methods, 12 evaluation tools, and automated evaluation pipelines.
Semantic-invariant watermarking for robust provenance of LLM-generated text.
Publicly verifiable watermarking for language models.
An open-source toolkit for generative watermarking of latent diffusion models.
04Earlier work: text-to-SQL & NLP robustness
THE JOURNEY
Research background.
EXPERIENCE
JUL 2025 — PRESENT
WeChat AI, Tencent
Researcher · WeLM Team
Core member of the WeLM pretraining team. Responsible for model architecture design; specializing in pretraining infrastructure.
Apple AIML
Research Intern · Advisor: Dr. Meng Cao
CUHK MISC Lab
Visiting Scholar · Advisor: Prof. Irwin King
UIC BDSC Lab
Visiting Scholar · Advisor: Prof. Philip S. Yu
EDUCATION
2020 — 2025
Tsinghua University
Ph.D. in Software Engineering
Advisor: Prof. Lijie Wen
Research in LLM alignment and
trustworthy LLMs.
2016 — 2020
Nanjing University
B.E. in Software Engineering
LET’S BUILD WHAT’S NEXT
Better models.
Bigger possibilities.
Open to research collaborations in foundation model architecture,
large-scale pretraining, and efficient generation.