Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Future Blog Post
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
Thought
Published:
I believe pretraining for large language models is far from finished. A lot of work today focuses on building agents and on post-training. Those steps can help on specific tasks, but the gains are often narrow. The deeper abilities we still need—making sense of the world and dealing with messy real situations—depend on large-scale pretraining with careful data and real changes to how the model is built. That includes new versions of the basic Transformer design, memory inside the model that ties ideas together more cleanly, and flexible ways to blend representations when the setting shifts.
Life in Shanghai 2025 Winter
Published:
I spent January and February 2026 interning at the LUMIA Lab at Shanghai Jiao Tong University. It was a busy time focused on ICML submissions, where I finished a first-author paper on dLLM and a third-author paper on PonderLM, among others. I chatted a lot with the seniors and my mentor in the group and learned a great deal.
Doing the obvious things right is not easy
Published:
Doing the obvious things right is harder than it sounds. I have watched “model companies” that barely build models. I have seen placeholder papers with no real experiments—work you cannot reproduce, or results that do not add up. Plenty of people train models by hoping for a lucky run. And I am sure many of us have seen bad ideas pushed forward anyway, because someone cared more about their own gain than about whether the thing actually works.
Internship in Beijing as A Quant in 2025 Summer
Published:
I applied for countless internships in Shanghai, but somehow, I ended up in Beijing. I still remember the excitement when I received the internship offer in early July. I thought I was blazing a new trail, so I came here, rented an apartment, and started my internship life.
DeltaNet Notes
Published:
论文记录《Parallelizing Linear Transformers with the Delta Rule over Sequence Length》
Blog Post number 4
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
publications
Detecting Violations of Physical Common Sense in Images: A Challenge Dataset and Effective Model
Published in ACM Multimedia 2025 (CCF-A), 2025
ACM Multimedia 2025 · First student author
Recommended citation: Weibin Wu, Zitong Wang, Zhengjie Luo, Wenqing Chen, Zibin Zheng. Detecting Violations of Physical Common Sense in Images. ACM Multimedia 2025.
Download Paper
Every Step a Thought: Implicit Visual Reasoning in Diffusion Language Models
Published in Under review, 2026
Under review
Recommended citation: Zitong Wang, Haohao Xu, Zijun Shen, Weibin Wu, Zibin Zheng. Every Step a Thought: Implicit Visual Reasoning in Diffusion Language Models. Under review.
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
Published in arXiv preprint, 2026
arXiv preprint
Recommended citation: Boyi Zeng, Yiqin Hao, He Li, Shixiang Song, Feichen Song, Zitong Wang, Siyuan Huang, Yi Xu, Ziwei He, Xinbing Wang, Zhouhan Lin. Pretraining with Token-Level Adaptive Latent Chain-of-Thought. arXiv preprint.
Download Paper
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
Published in arXiv preprint, 2026
arXiv preprint
Recommended citation: Shixiang Song, He Li, Zitong Wang, Boyi Zeng, Feichen Song, Yixuan Wang, Ziwei He, Zhouhan Lin. AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth. arXiv preprint.
Download Paper
Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation
Published in arXiv preprint, 2026
arXiv preprint
Recommended citation: Zitong Wang, Zijun Shen, Haohao Xu, Zhengjie Luo, Weibin Wu. Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation. arXiv preprint.
Download Paper
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
Published in ICML 2026 Spotlight (top 2.2%), 2026
ICML 2026 Spotlight (top 2.2%)
Recommended citation: Boyi Zeng, He Li, Shixiang Song, Yixuan Wang, Zitong Wang, Ziwei He, Xinbing Wang, Zhouhan Lin. PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space. ICML 2026 (Spotlight).
Download Paper
reading
Sample Poem or Quote
Poem · 2025
First line or short description.
Atomic Habits
Essay · 2026
One-sentence summary or why you recommend it.
Poor Charlie’s Almanack
Author Name · Book · 2026
talks
Talk 1 on Relevant Topic in Your Field
Published:
This is a description of your talk, which is a markdown file that can be all markdown-ified like any other post. Yay markdown!
Conference Proceeding talk 3 on Relevant Topic in Your Field
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Teaching experience 2
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.