Multimodal AI · Video understanding · Inference systems

Wayner Barrios

I build AI that can perceive and reason about the world more like we do.

My research spans multimodal perception, video understanding, reasoning, reinforcement-learning post-training, and efficient AI systems.

Portrait of Wayner Barrios
Dartmouth CollegePh.D.

01 / About

Beyondrecognition.

I develop models that connect perception with higher-level understanding: following events through time, grounding language in visual evidence, integrating multiple modalities, and learning coherent reasoning strategies.

I work across video understanding, visual grounding, reinforcement-learning post-training, and transparent reasoning evaluation. Recent projects include CRYSTAL, a benchmark and training framework for evaluating intermediate reasoning steps; MoDA, for fine-grained grounding in instructional multimodal models; and learnable attention mechanisms for aligning information across modalities.

I also build high-performance AI infrastructure. I created vLLM-MLX, an OpenAI- and Anthropic-compatible inference server for Apple Silicon that reaches more than 400 tokens per second, and at Samsung Research America I worked on reinforcement-learning post-training for multimodal video reasoning.

My work has appeared at CVPR, ICCV, ECCV, ICML, WACV, and other leading venues, as well as in PLOS Digital Health, the Journal of Pathology Informatics, Frontiers in Medical Technology, and other journals. I have collaborated with DARPA/IARPA, Adobe Research, Samsung Research, Meta AI, Carnegie Mellon University, KAUST, Northeastern University, Mount Sinai, Lunenfeld Institute, Universidad del Norte, EAFIT, Universidad CES, Universidad de Antioquia, among others.

I hold a Ph.D. in Computer Science from Dartmouth College, advised by SouYoung Jin, and founded Wiqonn to grow ambitious AI work in Latin America.

02 / Selected research

Perception, learning,and reasoning.

Earlier publications 4 papers
WACVW2019

Minding the Gaps in a Video Action Analysis Pipeline

J. Chen, J. Liu, J. Liang, T. Y. Hu, W. Ke, Wayner Barrios, D. Huang, A. G. Hauptmann

Analyzes the training and testing gaps between proposal, classification, and localization modules in event detection pipelines, then introduces practical fixes.

Complete publication recordGoogle Scholar

03 / Open source

Research thatruns.

Projects I build, plus contributions to major open-source communities — inference, evaluation, and geospatial infrastructure.

01vLLM-MLXHigh-throughput, OpenAI- and Anthropic-compatible LLM and MLLM inference for Apple Silicon, with continuous batching, tool calling, vision, audio, and video support.1.5k+stars400+tok/s 02GPT4All ContributionMerged a dataset-handling fix into GPT4All's local LLM runtime (77k+ stars) — the most widely adopted on-device LLM stack.77kstars1merged PR 03GeoNode ContributionGeoNode, the leading open-source geospatial platform. I contributed Docker/GeoServer OAuth2 integration and the notification refactor (3 merged PRs), plus the Dockerized project skeleton (33 commits) and fixes across the GeoNode Docker stack.1.7kstars3merged PRs 04OpenCode Power PackProduction workflows for code review, security, feature development, frontend design, and project memory, ported from Claude Code to OpenCode.11skills450+stars 05MapProxy ContributionAdded dimension-layer caching for WMS and WMTS to MapProxy — the tile cache and WMS proxy at the core of OpenStreetMap infrastructure.666stars10commits 06CRYSTALECCV 2026 benchmark, dataset, and pip-installable metrics for evaluating ordered reasoning traces and training with process rewards.6,372instancespippackage 07MoDAOfficial implementation of MoDA (ICML 2026) — instruction-guided channel modulation for fine-grained visual grounding in instructional MLLMs, with under 1% additional FLOPs.ICML2026<1%extra FLOPs 08Localizing MomentsOfficial PyTorch implementation of "Localizing Moments in Long Video Via Multimodal Guidance" (ICCV 2023) — identifies describable windows and prunes irrelevant segments before matching a query.ICCV202323stars 09Hypermap Registry ContributionElasticSearch API, aggregations, and time faceting for Harvard CGA's Hypermap Registry — remote map services made easy for spatial data infrastructures.15PRs9merged 10WorldMap ContributionHarvard CGA's WorldMap. I maintained the GeoNode 2.6-based platform: Django migrations, Dataverse/DataTables integration, and Dockerized deployment — 66 commits and 6 merged PRs.66commits6PRs merged 11ActivityNet ContributionBuilt and maintained the official ActivityNet challenge website — evaluation server, leaderboard, and GPU awards — for the leading video understanding benchmark.114commits2016challenge site 12DGX Spark Fine-tune LLMAn experimental end-to-end pipeline for LoRA fine-tuning, NVFP4/MXFP8 quantization, TensorRT-LLM export, and OpenAI-compatible serving on Blackwell GB10.4/8bit trainingGB10Blackwell

Open to research scientist and research engineering roles in multimodal AI, video understanding, post-training, and inference systems.

Get in touch