The query encoder does not need to see. A 70M text-only student keeps 95% of a 2B teacher's quality at 50x lower CPU latency.
News
- Sep 2026 DistilVDR is accepted to DocInsights 2026, the Workshop on Document Intelligence and Understanding at EMNLP 2026.
- Aug 2026 Two papers, NanoVDR and SAP, are accepted to EMNLP 2026 (Main). See you in Budapest!
- Jul 2026 Giving an oral presentation at the ACM SIGCHI Symposium on Engineering Interactive Computing Systems (EICS 2026) in Greece on our human-AI collaboration for sensor placement in human motion sensing work.
- Apr 2026 Our paper on human-AI collaboration in LLM-assisted design is accepted to Proceedings of the ACM on Human-Computer Interaction.
Earlier news (4)
- Apr 2026 IKEA-Bench released — 1,623 questions benchmarking 19 VLMs on cross-depiction assembly alignment.
- Apr 2026 Academic visit to the University of Southampton, UK.
- Mar 2026 NanoVDR is now on HuggingFace — full datasets, models, and an online demo.
- Oct 2025 Academic visit to TUM and LMU Munich.
Visual Document Retrieval
Multi-vector distillation does not need documents at all. Optimal transport aligns student and teacher query tokens, keeping 95% of teacher quality at 26x faster encoding.
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
NanoVDR shrank only the query encoder. DistilVDR distills the whole retriever end-to-end from a single 8B teacher, keeping most of its quality at 524M with a 15x smaller index.
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
A document's structure lives in the middle layers, not the final one. Pruning by that signal removes over 90% of the index without any training.
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
Visual Document Understanding
With a language-cue interface, we find the grounding bottleneck is the coordinate output, not the model's ability to locate the evidence.
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Adding text recovers instruction understanding but breaks diagram-to-video alignment. Scaling up does not help, because the bottleneck is the visual encoder.
IKEA-Bench: Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment
Smart Wearable & Motion Understanding
LLM-assisted design does not help everyone equally. The least experienced designer improved through collaboration, the most experienced one regressed.
Exploring Human-AI Collaboration in E-Textile Design: A Case Study on Flex Sensor Placement for Shoulder Motion Detection
Human motion does not need an encoder. Describe it in structured words and let the LLM do the reasoning.
Encoder-Free Human Motion Understanding via Structured Motion Descriptions
A 3D skeleton becomes a pseudo-image, so motion can be retrieved like a document, with token-level explainability.
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
Human motion gets its own alphabet, words and grammar, so the representation stays interpretable and unambiguous.