I develop models that connect perception with higher-level understanding: following events through time, grounding language in visual evidence, integrating multiple modalities, and learning coherent reasoning strategies.
I work across video understanding, visual grounding, reinforcement-learning post-training, and transparent reasoning evaluation. Recent projects include CRYSTAL, a benchmark and training framework for evaluating intermediate reasoning steps; MoDA, for fine-grained grounding in instructional multimodal models; and learnable attention mechanisms for aligning information across modalities.
I also build high-performance AI infrastructure. I created vLLM-MLX, an OpenAI- and Anthropic-compatible inference server for Apple Silicon that reaches more than 400 tokens per second, and at Samsung Research America I worked on reinforcement-learning post-training for multimodal video reasoning.
My work has appeared at CVPR, ICCV, ECCV, ICML, WACV, and other leading venues, as well as in PLOS Digital Health, the Journal of Pathology Informatics, Frontiers in Medical Technology, and other journals. I have collaborated with DARPA/IARPA, Adobe Research, Samsung Research, Meta AI, Carnegie Mellon University, KAUST, Northeastern University, Mount Sinai, Lunenfeld Institute, Universidad del Norte, EAFIT, Universidad CES, Universidad de Antioquia, among others.
I hold a Ph.D. in Computer Science from Dartmouth College, advised by SouYoung Jin, and founded Wiqonn to grow ambitious AI work in Latin America.