Multimodal AI · Video understanding · Inference systems

Wayner Barrios Multimodal AI Researcher

I build AI that can perceive and reason about the world more like we do.

My research spans multimodal perception, video understanding, reasoning, reinforcement-learning post-training, and efficient AI systems.

Portrait of Wayner Barrios
Dartmouth CollegePh.D.

Research focus

From perception
to deployment.

Perceive
Multimodal AIVideo understanding
Reason
ReasoningPost-training
Build
Inference systemsOpen source

About

Bio

I'm Wayner Barrios, a multimodal AI researcher developing models that connect perception with higher-level understanding. My past work focused on video understanding and visual perception, and I'm now expanding into multimodal models that follow events through time, ground language in visual evidence, and learn coherent reasoning strategies.

I recently completed my Ph.D. in Computer Science at Dartmouth College. I work across video understanding, visual grounding, reinforcement-learning post-training, and transparent reasoning evaluation.

I'm also interested in efficient AI: building smaller, more capable models and the infrastructure and systems that make them practical. This includes high-performance inference systems—continuous batching, multimodal prefix caching, and native model serving—across Apple Silicon, edge devices, and other resource-constrained hardware.

My work has appeared at CVPR, ICCV, ECCV, ICML, WACV, and other leading venues, as well as in PLOS Digital Health, the Journal of Pathology Informatics, Frontiers in Medical Technology, and other journals.

The connection across this work is the path from visual evidence to an answer that can be inspected and a model that can be used. Grounding asks which parts of an image or video support a question. Reasoning asks whether the model follows those observations through to a coherent conclusion. Efficient inference asks how to make that process practical when computation and memory are limited.

These questions appear in different forms across my projects. MoDA studies how instructions shape visual features. CRYSTAL examines the intermediate steps in multimodal explanations. vLLM-MLX focuses on the systems that serve language and multimodal models on Apple Silicon. The project guides explain the methods, the reported results, and where to start with the code, so you can follow each idea from its research motivation to its implementation.

I earned my Ph.D. at Dartmouth College, advised by SouYoung Jin, and founded Wiqonn, an AI lab growing ambitious AI work in Latin America.

Selected research

Research thatruns.

All research publications11 papers

Open source

Built inthe open.

My research and tools first, then contributions to open-source communities — video, geospatial, and LLM infrastructure.

01

CRYSTAL

Benchmark data, evaluation metrics, and process-reward training code for the CRYSTAL multimodal reasoning framework.

6,372instancespippackage
02

MoDA

Code for instruction-guided channel modulation in multimodal language models, accompanying the MoDA paper.

ICML2026<1%extra FLOPs
All open source projects12 projects

Building Wiqonn and researching multimodal perception, video understanding, and efficient inference.

Get in touch