The 9th Annual Conference on Machine Learning and Systems (MLSys) took place in May in Bellevue, Washington. MLSys is a highly selective interdisciplinary conference sitting at the intersection of machine learning (ML) and systems design. The conference highlights cutting-edge research that combines generative AI, natural language processing, computer vision and reinforcement learning with infrastructure, deployment and hardware optimizations to make AI faster, scalable and more performant.
MLSys offered Capital One associates the opportunity to learn from world-class conference sessions presented by experts in the field. All the attending associates left brimming with new ideas and planned collaborations. Kel Vanee, MVP, Machine Learning Engineering, presented some of the work happening at Capital One on using AI to make AI more efficient.
Takeaways and favorite papers from MLSys 2026
Some of the most prevalent topics at MLSys this year were on cache management, model speculation, retrieval augmented generation (RAG) and agentic AI. With a plethora of relevant and interesting talks, we had no shortage of papers to choose favorites from. While a complete list of the papers we loved would be far too long, here are a few standouts:
Large language model inference optimization
One of the leading themes this year was how to more efficiently serve LLM models. We especially liked the papers on reducing self-attention costs, such as MAC-Attention: a Match–Amend–Complete scheme for fast and accurate attention computation and BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding. We found valuable insights in papers covering how best to overlap computation with communication, such as TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference, Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token and FlashAgents: Accelerating Multi-Agent LLM Systems via Streaming Prefill Overlap.
Retrieval augmented generation (RAG)
There were papers covering efficiencies on the whole RAG lifecycle. Since we rely on latency-constrained RAG applications at Capital One, we found special interest in papers covering recall optimization, such as LEANN: A Low-Storage Overhead Vector Index and When Enough is Enough: Rank-Aware Early Termination for Vector Search. We were also intrigued by the continued theme of overlapping communication with computation in RAG by lookahead prefetching of context during generation in TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval.
Agentic AI
The field is increasingly moving toward agentic AI, and the publications at MLSys reflected that fact with a strong emphasis on how to best design systems optimized for Agentic workflows. As we continue to evolve the use of agentic AI within Capital One, papers on optimizations for agent planning and memory management, such as AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents and Hippocampus: An Efficient and Scalable Memory Module for Agentic AI, present excellent insight into how to do so effectively. We also learned more about emerging trends in agentic deployment optimization with PROMPTS: PeRformance Optimization via Multi-Agent Planning for LLM Training and Serving where agents evaluate, diagnose and optimize LLM training and serving. Ultimately we found that mature agentic workloads have distinct patterns of use, allowing real wins to arise from co-designing the compute-stack and agents around each other.

Engagement and a hosted dinner at MLSys 2026
In addition to delivering and attending talks, Capital One associates at MLSys connected with graduate students and industry professionals in the Expo Floor booth, and extended many of these conversations at a networking event hosted at W Bellevue later in the week. At the event, Capital One associates engaged attendees in discussions about research and their personal experiences with systems design for ML. For those attendees evaluating new career choices, Capital One associates described the array of complex problems we are solving, as well as the foundational teams we are building to support both research and business-critical applied work.
Overall, MLSys 2026 was an excellent opportunity for Capital One associates to learn and share experiences with leading scientists from around the world. Returning from MLSys, our associates brought new perspectives, research ideas and inspiration to share with our teams and the broader Capital One knowledgebase as we continue to better understand how we can apply AI/ML to transform financial services.
Explore Capital One's AI research and career opportunities
Interested in joining a world-class team that is accelerating state-of-the-art AI research? Explore Applied Research jobs at Capital One.

Sean is an AI researcher with over a decade of experience focusing on ML architecture, optimization, and inference efficiency. He joined Capital One's Specialist Models - PRISM team after serving as the head of research at two West Coast startups and several years with the Federal Government. He leads technical development for Capital One's internal GenAI Agent Assist program.
Related blogs
Use left and right arrow buttons to navigate related blog post cards, or swipe on touch devices.
LLM reasoning and agentic safety at ICML 2026
Explore our latest research in critique-guided distillation and multi-turn agent uncertainty in Seoul.
Capital One Science | July 1, 2026
Capital One at ACL 2026
Discover how Capital One is advancing state-of-the-art AI/ML science through collaborative natural language processing research.
Capital One Science | June 30, 2026
Insights from the inaugural Capital One AI Symposium
Advancing the state of the art through multi-sector partnerships.
Capital One Science | April 23, 2026
