I'm a BA in Artificial Intelligence student at CUHK working on efficient multimodal models. I study which capabilities of vision–language models degrade, and how, when the visual tokens they see are compressed. I also write a set of notes on markets.
News
- Jun 2026Launched this site.
- 2026Working toward a workshop submission on visual-token compression.
Research
- Visual-token compression for VLMs — merging, pruning, pooling.
- Capability-specific degradation under a fixed token budget.
- Omission vs. hallucination — and what causes which.
More on the research page.
Projects
glassbox ↗
Build modern AI from scratch — LLM internals, RL post-training, and vision, one transparent mechanism per file.
frame2frame ↗
Head-pose estimation from video — Euler angles (yaw/pitch/roll) with FIR signal smoothing.
ItineralTrace ↗
A multimodal AI travel-planning agent with voice, vision, and memory.
Notes
Notes on markets — theses, position rationale, and post-mortems.
📬 Subscribe to get new investing notes by email on Substack.
- First note: how I'll use this spaceJun 29, 2026