Multimodal AI · Video understanding · Inference systems

Wayner Barrios

I build AI that can perceive and reason about the world more like we do.

My research spans multimodal perception, video understanding, reasoning, reinforcement-learning post-training, and efficient AI systems.

Portrait of Wayner Barrios
Dartmouth CollegePh.D. · 2026

01 / About

Beyondrecognition.

I develop models that connect perception with higher-level understanding: following events through time, grounding language in visual evidence, integrating multiple modalities, and learning coherent reasoning strategies.

I work across video understanding, visual grounding, reinforcement-learning post-training, and transparent reasoning evaluation. Recent projects include CRYSTAL, a benchmark and training framework for evaluating intermediate reasoning steps; MoDA, for fine-grained grounding in instructional multimodal models; and learnable attention mechanisms for aligning information across modalities.

I also build high-performance AI infrastructure. I created vLLM-MLX, an OpenAI- and Anthropic-compatible inference server for Apple Silicon that reaches more than 400 tokens per second, and at Samsung Research America I worked on reinforcement-learning post-training for multimodal video reasoning.

My work has appeared at CVPR, ICCV, ECCV, ICML, WACV, and other leading venues, as well as in PLOS Digital Health, the Journal of Pathology Informatics, Frontiers in Medical Technology, and other journals. I have collaborated with DARPA/IARPA, Adobe Research, Samsung Research, Meta AI, Carnegie Mellon University, KAUST, Northeastern University, Mount Sinai, Lunenfeld Institute, Universidad del Norte, EAFIT, Universidad CES, Universidad de Antioquia, among others.

I hold a Ph.D. in Computer Science from Dartmouth College, advised by SouYoung Jin, and founded Wiqonn to grow ambitious AI work in Latin America.

02 / Selected research

Perception, learning,and reasoning.

Earlier publications 4 papers
WACVW2019

Minding the Gaps in a Video Action Analysis Pipeline

J. Chen, J. Liu, J. Liang, T. Y. Hu, W. Ke, Wayner Barrios, D. Huang, A. G. Hauptmann

Analyzes the training and testing gaps between proposal, classification, and localization modules in event detection pipelines, then introduces practical fixes.

Complete publication recordGoogle Scholar

03 / Open source

Research thatruns.

Inference, evaluation, post-training, and tools for AI developers.

Open to research scientist and research engineering roles in multimodal AI, video understanding, post-training, and inference systems.

Get in touch