AI Engineer World's Fair 2025

Tuesday, 3 June 2025 – Thursday, 5 June 2025 Marriott Marquis, San Francisco, CA

245 speakers Meet the line-up →
Sessions
218
Speakers
245
Days
3
Rooms
31

218 of 218 sessions

Search and filter
  • Workshops Workshop

    ComfyUI

    Tue 3 Jun, 09:00 – 09:50·Salons 2-6: Workshops

    Quick introduction to ComfyUI and what's new followed by a QA session.

  • Workshops Workshop

    Advanced: Reinforcement Learning, Kernels, Reasoning, Quantization & Agents

    Tue 3 Jun, 09:00 – 12:00·Foothill C: Workshops

    Why is Reinforcement Learning (RL) suddenly everywhere, and is it truly effective? Have LLMs hit a plateau in terms of intelligence and capabilities, or is RL the breakthrough they…

  • Workshops Workshop

    Intro to GraphRAG

    Tue 3 Jun, 09:00 – 10:20·Golden Gate Ballroom C: Workshops

    Learn the foundations of GraphRAG, starting with knowledge graph construction and then common retrieval patterns.

  • Workshops Workshop

    A2A & MCP: Automating Business Processes with LLMs

    Tue 3 Jun, 09:00 – 10:20·Foothill G1&2: Workshops

    Ever wished your webhooks could think for themselves? Join us to discover how A2A agents can transform passive webhook endpoints into intelligent workflow processors. In this…

  • Workshops Workshop

    Introduction to LLM serving with SGLang

    Tue 3 Jun, 09:00 – 10:20·SOMA: Workshops

    Do you want to learn how to serve models like DeepSeek and Qwen with SOTA speeds on launch day? SGLang is an open-source fast serving framework for LLMs and VLMs that generates…

  • Workshops Workshop

    Beyond Benchmarks: Strategies for Evaluating LLMs in Production

    Tue 3 Jun, 09:00 – 10:20·Nobhill A&B: Workshops

    Accuracy scores and leaderboard metrics look impressive—but production-grade AI requires evals that reflect real-world performance, reliability, and user happiness. Traditional…

  • Workshops Workshop

    Useful General Intelligence

    Tue 3 Jun, 09:30 – 10:30·Golden Gate Ballroom B: Workshops

    We’re all hearing that AI agents will enable AGI, but they can’t yet reliably perform even basic computer tasks. It turns out that getting AI to click, type, and scroll is more…

  • Workshops Workshop

    Forget RAG Pipelines—Build Production-Ready AI Agents in 15 Minutes

    Tue 3 Jun, 09:45 – 10:45·Golden Gate Ballroom A: Workshops

    Want to take advantage of your data, but don't want to reinvent RAG infrastructure? Join our workshop and see how you can deploy Agentic RAG in minutes using Contextual AI's…

  • Workshops Workshop

    The AI Engineer’s Guide to Raising VC

    Tue 3 Jun, 09:55 – 10:20·Salons 2-6: Workshops

    A no fluff, all tactics discussion. More AI engineers should build startups, the world needs more software. But there’s a way to raise VC and it’s hard to do it if you’ve never…

  • Workshops Workshop

    Mastering AI Evaluation: From Playground to Production with Braintrust

    Tue 3 Jun, 10:40 – 12:00·Golden Gate Ballroom C: Workshops

    This hands-on workshop will guide participants through the complete AI evaluation lifecycle using Braintrust, from initial prompt testing to production monitoring. Attendees will…

  • Workshops Workshop

    Automating Escrow with USDC and AI

    Tue 3 Jun, 10:40 – 12:00·Golden Gate Ballroom B: Workshops

    This workshop explores how USDC, AI, and smart contracts can streamline escrow by automating fund release based on task or process verification. By using AI to interpret off-chain…

  • Workshops Workshop

    Solving for the hardest Eval challenge: Building Metrics that actually work

    Tue 3 Jun, 10:40 – 12:00·Foothill G1&2: Workshops

    One of the biggest challenges in building evals you can trust is building metrics that reliably measure goodness in your application; metrics that are highly accurate, rapid fast,…

  • Workshops Workshop

    Ship Agents that Ship: A Hands-On Workshop for SWE Agent Builders

    Tue 3 Jun, 10:40 – 12:00·Nobhill A&B: Workshops

    Coding agents are transforming how software gets built, tested, and deployed, but engineering teams face a critical challenge: how to embrace this automation wave without…

  • Microsoft Talk

    Piloting agents in GitHub Copilot

    Tue 3 Jun, 10:40 – 12:00·Nobhill C&D: Microsoft

    The agent capabilities added to GitHub Copilot have enhanced its ability to act as a peer programmer. Copilot can now discover and generate code based on existing standards, run…

  • Workshops Workshop

    Building Multimodal AI Agents (From Scratch)

    Tue 3 Jun, 10:40 – 12:00·SOMA: Workshops

    In this hands-on workshop, you will build a multimodal AI agent capable of processing mixed-media content—from analyzing charts and diagrams to extracting insights from documents…

  • Workshops Workshop

    Building Voice Agents with Gemini and Pipecat

    Tue 3 Jun, 11:00 – 12:00·Golden Gate Ballroom A: Workshops

    Voice AI Agents are being deployed today in a wide range of business contexts. For example: - handling an increasing variety of call center tasks, - collecting patient data…

  • Workshops Workshop

    Build multilingual Conversational AI Agents

    Tue 3 Jun, 11:15 – 12:15·Salons 2-6: Workshops

    In this workshop you will learn how to build multilingual Conversational AI agents that can automatically detect your user's spoken language and can seamlessly switch to their…

  • Workshops Workshop

    How LLMs work for Web Devs: GPT in 600 lines of Vanilla JS

    Tue 3 Jun, 12:00 – 13:00·Golden Gate Ballroom A: Workshops

    Don't be intimidated. Modern AI can feel like magic, but underneath the hood are principles that web developers can understand, even if you don't have a machine learning…

  • Workshops Workshop

    AI Engineering with the Google Gemini 2.5 Model Family

    Tue 3 Jun, 13:00 – 15:00·SOMA: Workshops

    Hands on Workshop on learning to use Gemini 2.5 Pro in combination with Agentic tooling and MCP Servers.

  • Workshops Workshop

    Information Retrieval from the Ground Up

    Tue 3 Jun, 13:00 – 15:00·Foothill G1&2: Workshops

    Vector search is only a feature. Search engines and information retrieval have retaken their position as the foundation of RAG. This workshop takes you through decades of research,…

  • Workshops Workshop

    From Mixture of Experts to Mixture of Agents … with Super Fast Inference

    Tue 3 Jun, 13:00 – 15:00·Nobhill A&B: Workshops

    Our hands-on workshop will walk you through how to build your own Mixture of Agents (MoA) system using the fastest, and most capable open models available: Qwen3-32B and Llama…

  • Workshops Workshop

    Real-World Development with GitHub Copilot and VS Code

    Tue 3 Jun, 13:00 – 15:00·Salons 2-6: Workshops

    Join us for a hands-on workshop designed to demonstrate how VS Code and GitHub Copilot's expanding suite of AI features can match or even surpasses the benefits of other popular AI…

  • Workshops Workshop

    Model-Maxxing: RFT, DPO, SFT (Fine-tuning with OpenAI)

    Tue 3 Jun, 13:00 – 15:00·Golden Gate Ballroom B: Workshops

    Covering all forms of fine-tuning and prompt engineering, like SFT, DPO, RFT, prompt engineering / optimization, and agent scaffolding.

  • Workshops Workshop

    Building Agents with Amazon Nova Act and MCP

    Tue 3 Jun, 13:00 – 15:00·Golden Gate Ballroom A: Workshops

    In this 2-hour workshop, participants will gain practical hands-on experience building sophisticated AI agents using Amazon's agent technologies. You'll learn to build agents that…

  • Workshops Workshop

    Graph Intelligence: Enhance Reasoning and Retrieval Using Graph Analytics

    Tue 3 Jun, 13:00 – 15:00·Golden Gate Ballroom C: Workshops

    Advanced GraphRAG techniques apply graph ML and algorithms, wrapped into tidy notebooks.

  • Workshops Workshop

    Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP

    Tue 3 Jun, 13:00 – 15:00·Foothill C: Workshops

    We've all seen website chat bots which can look up an order or answer a basic question -- but what does it take to build autonomous agents which manage long, delicate processes…

  • Workshops Workshop

    AI Red-Teaming and Prompt Engineering

    Tue 3 Jun, 15:30 – 17:30·SOMA: Workshops

    Learn about learnprompting.org and HackAPrompt, the first guide on prompt engineering in the world and the first competition on prompt injection respectively. I will give a…

  • Workshops Workshop

    AI Pipelines and Agents in Pure TypeScript with Mastra.ai

    Tue 3 Jun, 15:30 – 17:30·Salons 2-6: Workshops

    This hands-on workshop introduces Mastra.ai, a TypeScript framework that streamlines the development of agentic AI systems compared to traditional approaches using LangChain and…

    • Nick Nisi Software developer and panelist on the JS Party podcast
    • Zack Proser Open source hacker. Dev Education at WorkOS
  • Workshops Workshop

    Agentic Coding with Windsurf

    Tue 3 Jun, 15:30 – 17:30·Golden Gate Ballroom A: Workshops

    Agentic coding marks a new era in software development, where AI agents take on autonomous roles in coding tasks. The Windsurf IDE embodies this shift by integrating intelligent…

  • Workshops Workshop

    Building voice agents with OpenAI

    Tue 3 Jun, 15:30 – 17:30·Golden Gate Ballroom B: Workshops

    We'll walk through the differences between chained and speech-to-speech powered voice agents, how to approach them, best practices and transform a text-based agent into our first…

  • Workshops Workshop

    How to build world-class AI products (featuring Sarah Sachs, AI lead @ Notion)

    Tue 3 Jun, 15:30 – 17:30·Golden Gate Ballroom C: Workshops

    Join us for a hands-on workshop where you'll learn practical strategies to evaluate AI applications throughout their lifecycle—from initial testing of prompts to ongoing monitoring…

  • Workshops Workshop

    Navigating deep context in legacy code with Augment Agent

    Tue 3 Jun, 15:30 – 17:30·Foothill C: Workshops

    Attendees will learn to use an AI coding agent as a fast and intuitive part of navigating and working with complex, production-grade legacy code bases. We will drop directly into…

  • Workshops Workshop

    VoiceVision RAG - Integrating Visual Document Intelligence with Voice Response

    Tue 3 Jun, 15:30 – 17:30·Foothill G1&2: Workshops

    In this workshop we will explore the integration of Colpali, a cutting-edge Vision based Retrieval Model, with voice synthesis for next-generation RAG systems. We'll demonstrate…

  • Workshops Workshop

    Shipping AI That Works: An Evaluation Framework for PMs

    Tue 3 Jun, 15:30 – 17:30·Nobhill A&B: Workshops

    GenAI is reshaping the product landscape, creating huge opportunities (along with new expectations) for product managers. Yet while prompt engineering and model tuning get the…

  • Unassigned Keynote

    Spark to System: Building the Open Agentic Web

    Wed 4 Jun, 09:20 – 09:40·Keynote/General Session (Yerba Buena 7&8)

    AI builders no longer ask whether to use agents—but how many and how fast. In this kickoff keynote, Microsoft’s Asha Sharma shows what happens when natural language creation meets…

    • Asha Sharma CVP, Head of Product, Microsoft AI Platform
  • Unassigned Keynote

    Conviction Session

    Wed 4 Jun, 09:40 – 10:00·Keynote/General Session (Yerba Buena 7&8)

    Please Fill in

  • Unassigned Keynote

    Simon Willison Keynote

    Wed 4 Jun, 10:00 – 10:20·Keynote/General Session (Yerba Buena 7&8)

  • Expo Sessions Talk

    Realtime conversational video with Pipecat and Tavus

    Wed 4 Jun, 10:40 – 10:55·Juniper: Expo Sessions

    Tavus shipped the world's first realtime video avatar platform last year. Developers use Tavus' conversational video APIs to create education, social, and customer support agents.…

  • Expo Sessions Talk

    Architecting Agent Memory: Principles, Patterns, and Best Practices

    Wed 4 Jun, 10:40 – 10:55·Willow: Expo Sessions

    In the rapidly evolving landscape of agentic systems, memory management has emerged as a key pillar for building intelligent, context-aware AI Agents. Inspired by the complexity of…

  • Expo Sessions Talk

    AI-powered entomology: Lessons from millions of AI code reviews

    Wed 4 Jun, 10:40 – 10:55·Nobhill A&B: Expo Sessions

    This talk will explore insights from millions of automated code reviews, revealing trends in bugs, vulnerabilities, and code health that Graphite’s AI code review agent have…

  • Expo Sessions Talk

    Vibe Coding in Production

    Wed 4 Jun, 10:55 – 11:10·Juniper: Expo Sessions

    What's the role of vibe coding in a production-grade applications? Join Augment Code's Matt McClernan as he speaks to engineering leaders about building with AI.

  • Expo Sessions Talk

    What does Enterprise Ready MCP mean?

    Wed 4 Jun, 10:55 – 11:10·Willow: Expo Sessions

    Everyone is building MCP servers: from Slack integrations to personal data tools. They're good demos, but not ready to turn into production. So, what does it take to make MCP…

  • Expo Sessions Talk

    Mastering Engineering Flow with Windsurf

    Wed 4 Jun, 10:55 – 11:10·Nobhill A&B: Expo Sessions

    As experienced engineers, especially senior and staff engineers, our focus shifts towards complex problem-solving, architectural decisions, and mentoring. While AI tools promise…

  • Agent Reliability Talk

    AI Automation that actually works: $100M, messy data, zero surprises

    Wed 4 Jun, 11:15 – 11:35·Foothill C: Agent Reliability

    We will review the different kinds of automation use-cases, and the approach we used, that will drive over a $100M of expected annual impact by deploying AI for business critical…

  • Product Management Keynote

    [PM Keynote] Everything is ugly so go build something that isn't

    Wed 4 Jun, 11:15 – 11:35·Foothill G 1&2: Product Management

    We're in an awkward adolescent phase of AI product (design). But what if this chaotic moment is actually our greatest opportunity? Enter the rebuilding revolution. In this talk,…

    • Raiza Martin CEO & Co-Founder Huxe || Previously NotebookLM
  • Infrastructure Talk

    What every AI engineer needs to know about GPUs

    Wed 4 Jun, 11:15 – 11:35·Foothill F: Infrastructure

    Every programmer needs to know a few things about hardware, like processors, memory, and disks. Due to AI systems' extreme demand for mathematical processing power, AI engineers…

  • Voice Keynote

    [Voice Keynote] Your realtime AI is ngmi

    Wed 4 Jun, 11:15 – 11:35·Foothill E: Voice

    Sean DuBois of OpenAI and Pion, and Kwindla Hultman Kramer of Daily and Pipecat, will talk about why you have to design realtime AI systems from the network layer up. Most…

  • MCP Talk

    MCP Origins & RFS

    Wed 4 Jun, 11:15 – 11:35·Yerba Buena Ballroom Salons 7-8: MCP

    Learn more about the latest updates on MCP and get ideas for what startups to build.

  • GraphRAG Talk

    HybridRAG: A Fusion of Graph and Vector Retrieval to Enhance Data Interpretation

    Wed 4 Jun, 11:15 – 11:35·Golden Gate Ballroom B: GraphRAG

    Interpreting complex information from unstructured text data poses significant challenges to Large Language Models (LLM), with difficulties often arising from specialized…

  • LLM RecSys Keynote

    Recsys Keynote: Improving Recommendation Systems & Search in the Age of LLMs

    Wed 4 Jun, 11:15 – 11:35·Golden Gate Ballroom A: LLM RecSys

    Recommendation systems and search have long adopted advances in language modeling, from early adoption of Word2vec for embedding-based retrieval to the transformative impact of…

  • AI in the Fortune 500 Talk

    From Copilot to Colleague: Building Trustworthy Productivity Agents for High-Stakes Work

    Wed 4 Jun, 11:15 – 11:35·Golden Gate Ballroom C: AI in the Fortune 500

    This keynote will explore what it takes to move from basic generative assistants to fully agentic AI—systems that don’t just suggest but plan, act, and adapt—all within the…

  • AI Architects Talk

    Rise of the AI Architect

    Wed 4 Jun, 11:15 – 11:35·SOMA: AI Architects

    As the amount of consumer facing AI products grows, the most forward leaning enterprises have created a new role: the AI Architect. These leaders are responsible for helping…

  • Agent Reliability Talk

    12 Factor Agents - Principles of Reliable LLM Applications

    Wed 4 Jun, 11:35 – 11:55·Foothill C: Agent Reliability

    Hi, I'm Dex. I've been hacking on AI agents for a while. I've tried every agent framework out there, from the plug-and-play crew/langchains to the "minimalist" smolagents of the…

  • MCP Talk

    What we learned from shipping remote MCP support at Anthropic

    Wed 4 Jun, 11:35 – 11:55·Yerba Buena Ballroom Salons 7-8: MCP

    We recently released remote MCP support for both claude.ai and the Anthropic API. This talk will cover architectural decisions we made in our implementation, remote MCP…

    • John Welsh Member of technical staff, Anthropic
  • Tiny Teams Talk

    The New Lean Startup

    Wed 4 Jun, 11:35 – 11:55·Yerba Buena Ballroom Salons 2-6: Tiny Teams

    In this session, I will be presenting a case study of Oleve's journey, revealing how we've scaled a profitable multi-product portfolio with a tiny team. I'll walk you through the…

  • LLM RecSys Talk

    What We Learned from Using LLMs in Pinterest Search

    Wed 4 Jun, 11:35 – 11:55·Golden Gate Ballroom A: LLM RecSys

    Pinterest Search integrates Large Language Models (LLMs) to enhance relevance scoring by combining search queries with rich multimodal content, including visual captions,…

  • GraphRAG Talk

    Wisdom Discovery at Scale: Code Less KAG with n8n MultiAI Agents

    Wed 4 Jun, 11:35 – 11:55·Golden Gate Ballroom B: GraphRAG

    "Wisdom Discovery at Scale: Code Less KAG with n8n MultiAI Agents"

  • AI in the Fortune 500 Talk

    Accelerating Investment Operations: How BlackRock Builds Custom Knowledge Apps at Scale.

    Wed 4 Jun, 11:35 – 11:55·Golden Gate Ballroom C: AI in the Fortune 500

    Investment Operations teams are the backbone of asset and investment management firms. Their day-to-day work not only enables portfolio managers to respond swiftly to market events…

  • AI Architects Talk

    Building Applications with AI Agents

    Wed 4 Jun, 11:35 – 11:55·SOMA: AI Architects

    Generative AI has dramatically shortened the distance between ideas and implementation, enabling faster prototyping and deployment than ever before. But while language models can…

  • Product Management Talk

    Why your product needs an AI product manager, and why it should be you

    Wed 4 Jun, 11:35 – 11:55·Foothill G 1&2: Product Management

    So you've built another cool demo. Now what? You have hype, but not impact. You have kudos but no users. Ultimately you have a demo, but not a product. The unique uncertainty of…

  • Infrastructure Talk

    Why We Don’t Need More Data Centers

    Wed 4 Jun, 11:35 – 11:55·Foothill F: Infrastructure

    AI infrastructure today is caught in an endless cycle: build more data centers, deploy more GPUs, repeat. But this approach is fundamentally flawed—expensive, inefficient, and…

  • Voice Talk

    Shipping an Enterprise Voice AI Agent in 100 Days

    Wed 4 Jun, 11:35 – 11:55·Foothill E: Voice

    What does it take to go from blank page to live enterprise voice agent in 100 days? That’s the challenge we took on with Fin Voice at Intercom. Enterprise customer service…

  • Tiny Teams Talk

    Using OSS models to build AI apps with millions of users

    Wed 4 Jun, 11:55 – 12:15·Yerba Buena Ballroom Salons 2-6: Tiny Teams

    In this talk, Hassan will go over how he builds open source AI apps that get millions of users like roomGPT.io (2.9 million users), restorePhotos.io (1.1 million users),…

  • LLM RecSys Talk

    360Brew LLM-based Foundation Model for Personalized Ranking and Recommendation

    Wed 4 Jun, 11:55 – 12:15·Golden Gate Ballroom A: LLM RecSys

    We will give a talk about our journey of building a foundation model for solving ranking and recommendation tasks across LinkedIn platform

  • AI in the Fortune 500 Talk

    How agents will unlock the $500B promise of AI

    Wed 4 Jun, 11:55 – 12:15·Golden Gate Ballroom C: AI in the Fortune 500

    AI agents are on the cusp of revolutionizing work as we know it. The number of use cases software can tackle is set to explode as AI handles tasks requiring real judgment. But to…

  • AI Architects Talk

    The Rise of Open Models in the Enterprise

    Wed 4 Jun, 11:55 – 12:15·SOMA: AI Architects

    This year kicked off with the DeepSeek-R1 news cycle breaking out of our AI Engineering bubble into the mainstream tech and business world. Leaders at the highest levels of the…

  • Agent Reliability Talk

    Scaling AI agents without breaking reliability

    Wed 4 Jun, 11:55 – 12:15·Foothill C: Agent Reliability

    As AI agents move from prototypes to production, developers are running into new challenges with orchestration, failure handling, and infrastructure. This session will unpack…

  • Product Management Talk

    Shipping something to someone always wins

    Wed 4 Jun, 11:55 – 12:15·Foothill G 1&2: Product Management

    Learnings from building products at Stripe and applying them in an AI native word

    • Kenneth Auchenberg Partner at AlleyCorp l, ex Product Lead at Stripe, Microsoft (VS Code)
  • Infrastructure Talk

    Large Scale AI on Apple Silicon using EXO

    Wed 4 Jun, 11:55 – 12:15·Foothill F: Infrastructure

    The hardware lottery: when a research idea wins because it is better suited to current hardware and software, and not because it is universally superior. Machine learning…

  • Voice Talk

    What we can learn from self driving in autonomous voice agents

    Wed 4 Jun, 11:55 – 12:15·Foothill E: Voice

    The reliability challenges facing voice & chat AI deployment today mirror those that the autonomous vehicle industry confronted years ago. This talk explores how evaluation…

  • GraphRAG Talk

    When Vectors Break Down: Graph-Based RAG for Dense Enterprise Knowledge

    Wed 4 Jun, 11:55 – 12:15·Golden Gate Ballroom B: GraphRAG

    Enterprise knowledge bases are filled with "dense mapping," thousands of documents where similar terms appear repeatedly, causing traditional vector retrieval to return the wrong…

  • MCP Talk

    Full Spectrum MCP: Uncovering Hidden Servers and Clients Capabilities

    Wed 4 Jun, 11:55 – 12:15·Yerba Buena Ballroom Salons 7-8: MCP

    The true power of Model Context Protocol emerges when clients and servers collaborate across the full spectrum of the specification. This talk presents practical examples of how VS…

  • GraphRAG Talk

    Stop Using RAG as Memory

    Wed 4 Jun, 12:15 – 12:35·Golden Gate Ballroom B: GraphRAG

    RAG is great for static knowledge retrieval—but terrible at memory. Vectorstore-based systems sold as memory lack relational and temporal awareness, leading agents astray with…

  • AI in the Fortune 500 Talk

    3 ingredients for building reliable enterprise agents

    Wed 4 Jun, 12:15 – 12:35·Golden Gate Ballroom C: AI in the Fortune 500

    It's easy to build a prototype of an agent, but hard to put an agent in production - especially in an enterprise setting. In this section, will talk about three ingredients for…

  • Agent Reliability Talk

    Production software keeps breaking and it will only get worse. Here’s how Traversal is fixing it.

    Wed 4 Jun, 12:15 – 12:35·Foothill C: Agent Reliability

    Software is eating the world. AI is eating software. AI-powered SWE means a whole lot more software is going to be written that powers mission critical systems in the coming years,…

  • Product Management Talk

    Survive the AI Knife-Fight: Building Products That Win

    Wed 4 Jun, 12:15 – 12:35·Foothill G 1&2: Product Management

    If you’ve ever been blocked by vague specs, shifting goals, or chasing “vibes,” things have only gotten messier in the age of AI. Everyone is obsessing over engineers doing PM work…

  • Voice Talk

    Serving Voice AI at $1/hr: Open-source, LoRAs, Latency, Load Balancing

    Wed 4 Jun, 12:15 – 12:35·Foothill E: Voice

    This is a talk that goes over our experience deploying Orpheus (Emotive, Realtime TTS) to production. It will cover topics: - Latency and optimizations - High fidelity voice…

  • MCP Talk

    MCP isn’t good, yet.

    Wed 4 Jun, 12:15 – 12:35·Yerba Buena Ballroom Salons 7-8: MCP

    You’ve heard a lot about MCP, probably been given an AI mandate or two, and are trying to figure out what’s real and what’s make believe. This session will give practical…

  • Tiny Teams Talk

    Gumloop's Path to be a 10 person unicorn

    Wed 4 Jun, 12:15 – 12:35·Yerba Buena Ballroom Salons 2-6: Tiny Teams

    An overview of how Gumloop is scaling automation across companies like Instacart, Webflow and Shopify with less than 10 people.

  • Microsoft Talk

    AI Red Teaming Agent: Accelerate your AI safety and security journey with Azure AI Foundry

    Wed 4 Jun, 12:45 – 13:05·Nobhill C&D: Microsoft

    In the age of autonomous AI agents, ensuring their safety and reliability is paramount. But how can we proactively uncover vulnerabilities before they impact real-world scenarios?…

  • Expo Sessions Talk

    How fast are LLM inference engines anyway?

    Wed 4 Jun, 12:45 – 13:00·Juniper: Expo Sessions

    Open weights models and open source inference servers have made massive strides in the year since we last got together at AIE World's Fair. Where once we had only pirated LLaMA…

  • Expo Sessions Talk

    How to trust an agent with software delivery

    Wed 4 Jun, 12:45 – 13:00·Willow: Expo Sessions

    AI-powered agents promise faster, easier software delivery, but their unpredictable behavior often makes engineers hesitant to fully trust them with critical workflows. Sam Alba,…

  • Expo Sessions Talk

    Events are the Wrong Abstraction for Your AI Agents

    Wed 4 Jun, 13:00 – 13:15·Willow: Expo Sessions

    AI Agents are distributed systems. Agents need to connect and communicate with tools, data repositories, other agents, etc., all over a network. Event-Driven Architecture is a…

  • Expo Sessions Talk

    Effective agent design patterns in production

    Wed 4 Jun, 13:00 – 13:15·Juniper: Expo Sessions

    At LlamaIndex we see a lot of agents built every day, and we've got a sense of what works and what doesn't. We've distilled those learnings down into a series of patterns and best…

  • Microsoft Talk

    Agentic Excellence: Mastering Evaluation of AI Agents with Azure AI Evaluation SDK

    Wed 4 Jun, 13:10 – 13:30·Nobhill C&D: Microsoft

    As AI agents transition from experimental assistants to critical components of enterprise workflows, reliably evaluating their performance becomes essential. But how do you…

  • Expo Sessions Talk

    Revenue Engineering: How to Price (and Reprice) Your AI Product

    Wed 4 Jun, 13:15 – 13:30·Juniper: Expo Sessions

    You’ve trained the model—now it’s time to train the business. This talk dives into the engineering behind pricing systems that can evolve as fast as your AI stack. Orb CTO…

  • Expo Sessions Talk

    Agentic GraphRAG: AI’s Logical Edge

    Wed 4 Jun, 13:30 – 13:45·Willow: Expo Sessions

    AI models are getting tasked to do increasingly complex and industry specific tasks where different retrieval approaches provide distinct advantages in accuracy, explainability,…

  • LLM RecSys Talk

    One model to rule recommendations: Netflix's Big Bet

    Wed 4 Jun, 14:00 – 14:20·Golden Gate Ballroom A: LLM RecSys

    Discuss the foundation model strategy for personalization at Netflix based on this post https://netflixtechblog.com/foundation-model-for-personalized-recommendation-1a0bd8e02d39…

    • Yesu Feng Netflix, Staff Research Scientist
  • GraphRAG Talk

    Practical GraphRAG - Making LLMs smarter with Knowledge Graphs

    Wed 4 Jun, 14:00 – 14:20·Golden Gate Ballroom B: GraphRAG

    RAG has become one standard architecture component for GenAI applications to address hallucinations and integrate factual knowledge. While vector search over text is common,…

  • AI in the Fortune 500 Talk

    Building an Agentic Platform

    Wed 4 Jun, 14:00 – 14:20·Golden Gate Ballroom C: AI in the Fortune 500

    Explore the technical evolution of metadata extraction at Box and how it shaped the foundation of our AI platform. We’ll walk through our transition to an agentic-first design—why…

  • AI Architects Talk

    AI That Pays: Lessons from Revenue Cycle

    Wed 4 Jun, 14:00 – 14:20·SOMA: AI Architects

    While much of the AI innovation in healthcare has centered on clinical and patient-facing applications, Revenue Cycle Management (RCM) remains an underexplored yet critical domain.…

  • Agent Reliability Talk

    How to build Enterprise-aware agents

    Wed 4 Jun, 14:00 – 14:20·Foothill C: Agent Reliability

    While LLMs demonstrated impressive reasoning capabilities, their out-of-the-box reasoning is akin to hiring a brilliant but brand-new employee who doesn’t have the enterprise…

  • Product Management Talk

    Shipping Products When You Don’t Know What they Can Do

    Wed 4 Jun, 14:00 – 14:20·Foothill G 1&2: Product Management

    A customer recently asked me: “Hey, can I tag your AI agent in a Google Doc comment?” The honest answer: I have no idea! We never designed our agents to handle Google Doc…

  • Infrastructure Keynote

    [Infra Keynote] Geopolitics of AI Infrastructure

    Wed 4 Jun, 14:00 – 14:20·Foothill F: Infrastructure

    As AI reshapes the global balance of power, the infrastructure behind it—chips, data centers, power, and supply chains—has become a new arena for geopolitical competition. This…

  • Voice Talk

    Building the Voice-First Future: Omnipresent Agents that Listen, Talk and Act

    Wed 4 Jun, 14:00 – 14:20·Foothill E: Voice

    We’re entering a world where talking to machines feels as natural as talking to people. Voice is about to become the dominant interface for technology - ambient, always-on, and…

  • Tiny Teams Talk

    Tiny Teams

    Wed 4 Jun, 14:00 – 14:20·Yerba Buena Ballroom Salons 2-6: Tiny Teams

    Sean reached out on X, happy to do a talk on how to build a tiny team

  • MCP Talk

    MCP is all you need

    Wed 4 Jun, 14:00 – 14:20·Yerba Buena Ballroom Salons 7-8: MCP

    Everyone is talking about agents, and right after that, they’re talking about agent-to-agent communications. Not surprisingly, various nascent, competing protocols are popping up…

  • LLM RecSys Talk

    Instacart’s LLM-driven approach to Search and Discovery

    Wed 4 Jun, 14:20 – 14:40·Golden Gate Ballroom A: LLM RecSys

  • GraphRAG Talk

    Multi-Agent AI and Network Knowledge Graphs for Change Management and Network Testing

    Wed 4 Jun, 14:20 – 14:40·Golden Gate Ballroom B: GraphRAG

    Traditional ticketing and testing workflows for change management and network operations often operate independently and lack critical real-world context and adaptive decision…

  • AI in the Fortune 500 Talk

    The Billable Hour is Dead; Long Live the Billable Hour?

    Wed 4 Jun, 14:20 – 14:40·Golden Gate Ballroom C: AI in the Fortune 500

    If software was eating the world before, knowledge work will soon be devoured by AI. In corporate America there are thousands of hours spent on rote tasks every day by employees,…

  • AI Architects Talk

    Structuring a modern AI team

    Wed 4 Jun, 14:20 – 14:40·SOMA: AI Architects

    You've been given an AI mandate but don't have additional headcount, what next? Re-skilling, up-skilling and team augmentation become essential to delivering on a new mandate. In…

  • Agent Reliability Talk

    Agents vs Workflows: Why Not Both?

    Wed 4 Jun, 14:20 – 14:40·Foothill C: Agent Reliability

    One current hot debate is should you make your top-level abstraction a ReAct type agent running in a loop? or should you make it a structured workflow graph? OpenAI is launching…

  • Product Management Talk

    Make your LLM app a Domain Expert: How to Build an LLM-Native Expert System

    Wed 4 Jun, 14:20 – 14:40·Foothill G 1&2: Product Management

    Vertical AI is a multi-trillion-dollar opportunity. But you can't build a domain-expert application simply by grabbing the latest LLMs off-the-shelf: you need a system for…

  • Infrastructure Talk

    Hacking the Inference Pareto Frontier for Cheaper and Faster Tokens Without Breaking SLAs

    Wed 4 Jun, 14:20 – 14:40·Foothill F: Infrastructure

    Your model works! It aces the evals! It even passes the vibe check! All that’s required is inference, right? Oops, you’ve just stepped into a minefield: -Not low-latency enough?…

  • Voice Talk

    Why ChatGPT Keeps Interrupting You

    Wed 4 Jun, 14:20 – 14:40·Foothill E: Voice

    ChatGPT Advanced Voice Mode isn’t interrupting just you. Interruptions, and turn-taking in general, are unsolved problems for all Voice AI agents. Nobody likes being cut short –…

  • MCP Talk

    Observable tools - the state of MCP observability

    Wed 4 Jun, 14:20 – 14:40·Yerba Buena Ballroom Salons 7-8: MCP

    AI Engineers deserve observable tools! MCP getting adoption means that less and less of your agents code is running under your control, and this has DX and observability…

  • Tiny Teams Talk

    Building Small AI Teams with Huge Impact

    Wed 4 Jun, 14:20 – 14:40·Yerba Buena Ballroom Salons 2-6: Tiny Teams

    tbd

  • AI in the Fortune 500 Talk

    Build Dynamic Products, and Stop the AI Sideshow

    Wed 4 Jun, 14:40 – 15:00·Golden Gate Ballroom C: AI in the Fortune 500

    AI across product, GTM, and strategy was a great approach in 2023, but by now, we all already know that AI is disrupting the global landscape and how business gets done. Now is the…

  • MCP Talk

    The rise of the agentic economy on the shoulders of MCP

    Wed 4 Jun, 14:40 – 15:00·Yerba Buena Ballroom Salons 7-8: MCP

    Thanks to MCP and all the MCP server directories, agents can now autonomously discover new tools and other agents. This lays down the foundation for the future agentic economy,…

  • Tiny Teams Talk

    Benchmarks Are Memes: How What We Measure Shapes AI—and Us

    Wed 4 Jun, 14:40 – 15:00·Yerba Buena Ballroom Salons 2-6: Tiny Teams

    Benchmarks shape more than just AI models—they shape our future. The things we choose to measure become self-fulfilling prophecies, guiding AI toward specific abilities and,…

  • LLM RecSys Talk

    Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+ DAU

    Wed 4 Jun, 14:40 – 15:00·Golden Gate Ballroom A: LLM RecSys

    YouTube recommendations drive the majority of video watch time for billions of daily users. Traditionally powered by large embedding models (LEMs), we're undertaking a fundamental…

  • GraphRAG Talk

    Beyond Documents: Implementing Knowledge Graphs in Legal Agents

    Wed 4 Jun, 14:40 – 15:00·Golden Gate Ballroom B: GraphRAG

    Structured Representations are pretty important in the law, where the relationships between clauses, documents, entities, and multiple parties matter. Structured Representation…

  • AI Architects Talk

    Building Effective Voice Agents

    Wed 4 Jun, 14:40 – 15:00·SOMA: AI Architects

    How to build production voice applications and learnings from working with customers along the way

  • Agent Reliability Talk

    Vibe Coding, with Confidence

    Wed 4 Jun, 14:40 – 15:00·Foothill C: Agent Reliability

    Everyone wants to do Vibe Code, even large Enterprises. But how can we ensure that the generated code is well-grounded with the dev team's code and software development standards?…

  • Product Management Talk

    Building the platform for agent coordination

    Wed 4 Jun, 14:40 – 15:00·Foothill G 1&2: Product Management

    Learn how we're evolving Linear into an operating system for engineering teams to ship product with agents as a first class citizen.

  • Infrastructure Talk

    Flipping the Inference Stack: Why GPUs Bottleneck Real-Time AI at Scale

    Wed 4 Jun, 14:40 – 15:00·Foothill F: Infrastructure

    AI inference today is stuck in a loop: throw more GPUs at the problem, scale horizontally, rinse and repeat. But that playbook is hitting a wall. Latency, cost, and energy grids…

  • Voice Talk

    Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat

    Wed 4 Jun, 14:40 – 15:00·Foothill E: Voice

    The Gemini Live API GA is now powered by Google's best cost-effective thinking model Gemini 2.5 Flash. We will do a deep dive on the capabilities that the Gemini Live API combined…

  • Expo Sessions Talk

    Taming Rogue AI Agents with Observability-Driven Evaluation

    Wed 4 Jun, 15:15 – 15:30·Nobhill A&B: Expo Sessions

    LLM agents often drift into failure when prompts, retrieval, external data, and policies interact in unpredictable ways. This session introduces a repeatable, metric-driven…

  • Expo Sessions Talk

    Vector Search Benchmark[eting]

    Wed 4 Jun, 15:15 – 15:30·Willow: Expo Sessions

    Every vector database out there is both faster and slower than any other competitor — if you believe all the benchmarketing out there. Let's turn the marketing into useful…

  • Expo Sessions Talk

    Data is Your Differentiator: Building Secure and Tailored AI Systems

    Wed 4 Jun, 15:30 – 15:45·Juniper: Expo Sessions

    As organizations seek to harness their proprietary data while maintaining security and compliance, Amazon Bedrock provides a comprehensive framework for building tailored AI…

  • Expo Sessions Talk

    Polar Signals Expo Session

    Wed 4 Jun, 15:30 – 15:45·Willow: Expo Sessions

  • Unassigned Keynote

    Windsurf everywhere, doing everything, all at once

    Wed 4 Jun, 16:25 – 16:45·Keynote/General Session (Yerba Buena 7&8)

    abstract tbd

  • Unassigned Keynote

    #define AI Engineer

    Wed 4 Jun, 16:45 – 17:25·Keynote/General Session (Yerba Buena 7&8)

    Greg Brockman's career and advice for AI Engineers

  • Unassigned Keynote

    A year of Gemini progress + what comes next

    Thu 5 Jun, 09:05 – 09:25·Keynote/General Session (Yerba Buena 7&8)

    Over the last year, Google and Gemini models have shown rapid progress across all dimensions (model, product, etc). Let's highlight all the work that has happened, how we got the…

  • Unassigned Keynote

    Thinking Deeper in Gemini

    Thu 5 Jun, 09:25 – 09:45·Keynote/General Session (Yerba Buena 7&8)

    Progress towards general intelligence has been marked by identifying fundamental intelligence bottlenecks within existing models and developing solutions that improve the…

    • Jack Rae Principal Research Scientist
  • Unassigned Keynote

    Why should anyone care about Evals?

    Thu 5 Jun, 09:45 – 09:50·Keynote/General Session (Yerba Buena 7&8)

    An introduction to the evals track

  • Unassigned Keynote

    Containing Agent Chaos

    Thu 5 Jun, 09:50 – 10:10·Keynote/General Session (Yerba Buena 7&8)

    AI agents promise breakthroughs but often deliver operational chaos. Building reliable, deployable systems with unpredictable LLMs feels like wrestling fog – testing outputs alone…

  • Unassigned Keynote

    The infrastructure for the singularity

    Thu 5 Jun, 10:10 – 10:30·Keynote/General Session (Yerba Buena 7&8)

    We're at an inflection point where AI agents are transitioning from experimental tools to practical coworkers. This new world will demand new infrastructure for RL training,…

  • Expo Sessions Talk

    Introducing Strands Agents, an Open Source AI Agents SDK

    Thu 5 Jun, 10:45 – 11:00·Nobhill A&B: Expo Sessions

    Building AI agents used to require complex orchestration, extensive scaffolding, and months of tuning. With Strands Agents, an open source SDK from AWS. You can now build, test,…

  • Expo Sessions Talk

    The fastest software dev workflow in the world: AI meets stacked diffs

    Thu 5 Jun, 10:45 – 11:00·Willow: Expo Sessions

    Learn the secrets behind the workflows that engineers at the fastest moving companies in the world are using to build software for billions of users worldwide. This workshop will…

  • Expo Sessions Talk

    Why Your Agent’s Brain Needs a Playbook: Practical Wins from Using Ontologies

    Thu 5 Jun, 10:45 – 11:00·Juniper: Expo Sessions

    You're trying to guide how your agents think and act. Code-orchestrated workflows are too rigid, but LLMs charting their own course feel too chaotic. When you need a middle ground,…

  • Expo Sessions Talk

    The State of AI-Powered Search and Retrieval

    Thu 5 Jun, 11:00 – 11:15·Juniper: Expo Sessions

    In this talk, we examine the state-of-the-art in AI-powered search and retrieval. We detail techniques for enhancing performance beyond base embedding models, including hybrid…

  • Expo Sessions Talk

    The Build-Operate Divide: Bridging Product Vision and AI Operational Reality

    Thu 5 Jun, 11:00 – 11:15·Willow: Expo Sessions

    Product leaders see AI possibilities. Operations teams see implementation chaos. That disconnect can kill promising AI features before they ever reach users. In this session,…

  • Expo Sessions Talk

    Pipecat Cloud: Enterprise Voice Agents Built On Open Source

    Thu 5 Jun, 11:00 – 11:15·Nobhill A&B: Expo Sessions

    Voice AI agents today can conduct natural, human-like conversations and perform a wide variety of tasks: customer support, lead qualification, healthcare patient intake, market…

  • Reasoning + RL Talk

    Training Agentic Reasoners

    Thu 5 Jun, 11:15 – 11:35·Yerba Buena Ballroom 2-6: Reasoning + RL

    This talk will be a technical deep dive into RL for agentic reasoning via multi-turn tool calling, similar to OpenAI's o3 and Deep Research. In particular, we'll cover: - When,…

  • AI in the Fortune 500 Talk

    From Hype to Habit: How We’re Building an AI-First SaaS Company—While Still Shipping the Roadmap

    Thu 5 Jun, 11:15 – 11:35·Golden Gate Ballroom C: AI in the Fortune 500

    What does it really take to move a modern SaaS company from AI experimentation to becoming truly AI-first? At Sprout Social, we’re in the midst of that…

  • AI Architects Talk

    Monetizing AI: From Zero to Profit

    Thu 5 Jun, 11:15 – 11:35·SOMA: AI Architects

    As AI continues to transform industries, companies are faced with the critical challenge of effectively monetizing AI-driven products in a way that captures value, ensures customer…

  • SWE Agents Talk

    Devin 2.0 and the Future of SWE

    Thu 5 Jun, 11:15 – 11:35·Yerba Buena Ballroom 7&8: SWE Agents

    A talk on the future of software engineering with Scott Wu of Cognition AI, the makers of Devin.

  • Retrieval + Search Talk

    Building AI Agents that actually automate Knowledge Work

    Thu 5 Jun, 11:15 – 11:35·Golden Gate Ballroom A: Retrieval + Search

    Agents are all the rage in 2025, and every single b2b SaaS startup/incumbent promises AI agents that can "automate work" in some way. But how do you actually build this? The…

  • Evals Talk

    On Engineering AI Systems that Endure The Bitter Lesson

    Thu 5 Jun, 11:15 – 11:35·Golden Gate Ballroom B: Evals

    Will discuss the principles for building AI software that underpin DSPy, highlighting the differences between conventional prompting (or finetuning/RL) versus the design and…

  • Security Talk

    Safety and security for code-executing agents

    Thu 5 Jun, 11:15 – 11:35·Foothill C: Security

    Code is the lingua franca for both software engineers and highly capable AI models. As we give agents the ability to build, test, and run code that they generate, the command line…

  • Design Engineering Talk

    UX Design Principles for (Semi) Autonomous Multi-Agent Systems

    Thu 5 Jun, 11:15 – 11:35·Foothill G 1&2: Design Engineering

    Autonomous or semi-autonomous multi-agent systems (MAS) involve exponentially complex configurations (system config, agent configs, task management and delegation, etc.). These…

  • Generative Media Talk

    The State of Generative Media Today

    Thu 5 Jun, 11:15 – 11:35·Foothill F: Generative Media

    Generative AI is reshaping the creative landscape, enabling the production of images, audio, and video with unprecedented speed and sophistication. This session offers an in-depth…

  • Autonomy + Robotics Talk

    Robotics: why now?

    Thu 5 Jun, 11:15 – 11:35·Foothill E: Autonomy + Robotics

    Sharing recent progress from Physical Intelligence and why it is an exciting time to push the frontier in general purpose robotics

  • SWE Agents Talk

    Your Coding Agent Just Got Cloned And Your Brain Isn't Ready

    Thu 5 Jun, 11:35 – 11:55·Yerba Buena Ballroom 7&8: SWE Agents

    Will the future engineer code alongside a single coding agent, or will they spend their day orchestrating many agents? Traditional development rewards synchronous focus. This…

  • Reasoning + RL Talk

    Measuring AGI: Interactive Reasoning Benchmarks

    Thu 5 Jun, 11:35 – 11:55·Yerba Buena Ballroom 2-6: Reasoning + RL

    ARC Prize Foundation is building the North Star for AGI—rigorous, open benchmarks that track reasoning progress in modern AI. We'll show why static AGI evaluations are useful, but…

  • Retrieval + Search Talk

    Scaling Enterprise-Grade RAG Systems: Lessons from the Legal Frontier

    Thu 5 Jun, 11:35 – 11:55·Golden Gate Ballroom A: Retrieval + Search

    In domains like law, compliance, and tax, building enterprise-grade RAG means very large scale, spikey workloads, a focus on accuracy, and non-negotiable privacy. In this talk,…

  • Evals Talk

    Turning Fails into Features: Zapier’s Hard-Won Eval Lessons

    Thu 5 Jun, 11:35 – 11:55·Golden Gate Ballroom B: Evals

    Every agent failure can be a roadmap to your next breakthrough. This talk reveals how Zapier's evaluation system transforms frustrating user experiences into targeted improvements,…

  • Security Talk

    The Unofficial Guide to Apple’s Private Cloud Compute

    Thu 5 Jun, 11:35 – 11:55·Foothill C: Security

    In October 2024, Apple released a new private AI technology onto millions of devices called “Private Cloud Compute”. It brings the same level of privacy and security a local device…

  • Design Engineering Talk

    Good design hasn’t changed with AI

    Thu 5 Jun, 11:35 – 11:55·Foothill G 1&2: Design Engineering

    Bad designs are still bad. AI doesn’t make it good. The novelty of AI makes the bad things tolerable, for a short time. Building great designs and experiences with AI have the same…

  • Generative Media Talk

    Veo 3 for developers

    Thu 5 Jun, 11:35 – 11:55·Foothill F: Generative Media

    This talk will briefly trace the history of video generation models before diving into Veo 3, Google DeepMind's latest state-of-the-art model that marks a significant leap by…

    • Paige Bailey Engineering Lead - Developer Relations @ Google DeepMind
  • Autonomy + Robotics Talk

    Real-time Experiments with an AI Co-Scientist

    Thu 5 Jun, 11:35 – 11:55·Foothill E: Autonomy + Robotics

    The sheer volume of data and complexity of modern scientific challenges necessitate tools that go beyond mere analysis. The vision of an "AI Co-scientist" – a true collaborative…

  • AI Architects Talk

    CIOs and Industry Leaders: Do You Trust Your AI’s Inferences?

    Thu 5 Jun, 11:35 – 11:55·SOMA: AI Architects

    Enterprise AI adoption is accelerating, but with it comes a hard question: Do we trust the model’s decisions? In this 18-minute talk, I’ll explore the invisible risks behind…

  • AI in the Fortune 500 Talk

    Machines of Buying & Selling Grace

    Thu 5 Jun, 11:35 – 11:55·Golden Gate Ballroom C: AI in the Fortune 500

    How to go beyond browser automation to truly agentic commerce, where AI can buy, sell and negotiate on behalf of users and merchants.

  • Reasoning + RL Talk

    Post-Training Open Models with RL for Autonomous Coding

    Thu 5 Jun, 11:55 – 12:15·Yerba Buena Ballroom 2-6: Reasoning + RL

    The models and techniques to build fully autonomous coding agents - not just coding copilots - are already here. In this talk, former Google DeepMind staff research scientist, now…

  • Retrieval + Search Talk

    Evaluating AI Search: A Practical Framework for Augmented AI Systems

    Thu 5 Jun, 11:55 – 12:15·Golden Gate Ballroom A: Retrieval + Search

    AI search is becoming the front door to information, whether through Retrieval-Augmented Generation (RAG), Search-Augmented Generation (SAG), or custom agents that synthesize…

  • Evals Talk

    Evals Are Not Unit Tests

    Thu 5 Jun, 11:55 – 12:15·Golden Gate Ballroom B: Evals

    How to think about evaluating a non-deterministic system — and how to actually succeed at it.

  • Security Talk

    Fuzzing in the GenAI Era

    Thu 5 Jun, 11:55 – 12:15·Foothill C: Security

    "Evaluation" is one of those concepts that every AI practitioner vaguely knows is important, but few practitioners truly understand. Is "eval" the dataset for measuring the quality…

  • Design Engineering Talk

    The Bitter Layout or: How I Learned to Love the Model Picker

    Thu 5 Jun, 11:55 – 12:15·Foothill G 1&2: Design Engineering

    Are conversational interfaces the future or, as many designers have suggested, a lazy solution that is bottlenecking AI-HCI? Despite well-documented usability issues, the design of…

  • Generative Media Talk

    Magic Editor Under the Hood: Weaving Generative AI into a Billion-User App

    Thu 5 Jun, 11:55 – 12:15·Foothill F: Generative Media

    Go behind the scenes of Google Photos' Magic Editor. Explore the engineering feats required to integrate complex CV and cutting-edge generative AI models into a seamless mobile…

  • Autonomy + Robotics Talk

    What Is a Humanoid Foundation Model? An Introduction to GR00T N1

    Thu 5 Jun, 11:55 – 12:15·Foothill E: Autonomy + Robotics

    Foundation models don’t just write or draw anymore—they’re starting to move. GR00T N1 is NVIDIA’s open Vision-Language-Action (VLA) foundation model for humanoid robots. Built…

  • AI in the Fortune 500 Talk

    How Intuit uses LLMs to explain taxes to millions of taxpayers

    Thu 5 Jun, 11:55 – 12:15·Golden Gate Ballroom C: AI in the Fortune 500

    I will talk about how Intuit uses LLMs to explain tax situations to Turbotax users. Users want explanations of their tax situations - this drives confidence in the product. Over…

  • AI Architects Talk

    Does AI Actually Boost Developer Productivity? (Stanford / 100k Devs Study)

    Thu 5 Jun, 11:55 – 12:15·SOMA: AI Architects

    Forget vendor hype: Is AI actually boosting developer productivity, or just shifting bottlenecks? Stop guessing. Our study at Stanford cuts through the noise, analyzing…

  • Design Engineering Talk

    AI and Game Theory: A Case Study on NYT's Connections

    Thu 5 Jun, 12:15 – 12:35·Foothill G 1&2: Design Engineering

    This session will examine the interplay between human intuition and artificial intelligence in puzzle-solving, using the popular New York Times Connections game as a practical case…

  • AI in the Fortune 500 Talk

    Ship it! Building Production-Ready Agents

    Thu 5 Jun, 12:15 – 12:35·Golden Gate Ballroom C: AI in the Fortune 500

    Explore the practical challenges and solutions for deploying AI agents in real-world production environments. Through detailed technical analysis and practical examples, we'll…

  • AI Architects Talk

    AX is the only Experience that Matters

    Thu 5 Jun, 12:15 – 12:35·SOMA: AI Architects

    If you’re building devtools for humans, you’re building for the past. Already a quarter of Y Combinator’s latest batch used AI to write 95% or more of their code. AI agents are…

  • SWE Agents Talk

    Don’t get one-shotted: Leveraging AI to test, review, merge, and deploy code

    Thu 5 Jun, 12:15 – 12:35·Yerba Buena Ballroom 7&8: SWE Agents

    As AI tools like GitHub Copilot and ChatGPT help engineers generate code at an unprecedented rate, the “outer loop”—reviewing, testing, merging, and deploying—becomes more vital…

  • Reasoning + RL Talk

    OpenThinker - Unreasonably Effective Reasoning Distillation at Scale

    Thu 5 Jun, 12:15 – 12:35·Yerba Buena Ballroom 2-6: Reasoning + RL

    Peel back the curtain on state of the art model post-training through the story of OpenThinker, a SOTA small reasoning model (outperforming DeepSeek distill), built in the open.…

  • Retrieval + Search Talk

    RAG in 2025: State of the Art and the Road Forward

    Thu 5 Jun, 12:15 – 12:35·Golden Gate Ballroom A: Retrieval + Search

    The talk will have three parts 1.Roadmap debate: RAG vs. finetuning vs. long-context 2.RAG today: benefits, challenges, and current solutions 3.RAG tomorrow: AI models do more…

  • Security Talk

    Securing Agents with Open Standards

    Thu 5 Jun, 12:15 – 12:35·Foothill C: Security

    Shipping AI agents that are safe for production means solving some tough identity and authorization challenges that are not always obvious at the prototype stage. In practice, this…

  • Generative Media Talk

    Small, sharp tools in the era of generative AI

    Thu 5 Jun, 12:15 – 12:35·Foothill F: Generative Media

    Replicate makes it easy to run thousands of AI models in the cloud without any machine learning expertise. In this session, we'll show how to make those models available to LLMs as…

  • Expo Sessions Talk

    Agents, Access, and the Future of Machine Identity

    Thu 5 Jun, 12:45 – 13:00·Willow: Expo Sessions

    AI agents are calling APIs, submitting forms, and sending emails—but how do you control what they’re allowed to do? As agents act on behalf of users or organizations, traditional…

    • Nick Nisi Software developer and panelist on the JS Party podcast
  • Microsoft Talk

    Building Protected MCP Servers

    Thu 5 Jun, 12:45 – 13:05·Nobhill C&D: Microsoft

    Join us to see how VS Code and GitHub Copilot's expanding suite of AI features can match or even surpasses the benefits of other popular AI developer tools. We'll focus on…

  • Expo Sessions Talk

    Building CISO-approved agent fleet architecture

    Thu 5 Jun, 12:45 – 13:00·Juniper: Expo Sessions

    Security is the biggest blocker for agent orchestration adoption in regulated industries for SWE agents. Gitpod's agent orchestration went from an originally self-hosted kubernetes…

  • Expo Sessions Talk

    Conquering Agent Chaos

    Thu 5 Jun, 13:00 – 13:15·Willow: Expo Sessions

    Agent deployments can be dicey, especially at first. This session goes over all the things that cause headache with deployments from serverless issues to networking issues - and…

  • Expo Sessions Talk

    Serving Voice AI at Scale

    Thu 5 Jun, 13:15 – 13:30·Nobhill A&B: Expo Sessions

    Real-Time Voice AI applications demand the lowest possible latencies to enhance user experiences with more advanced reasoning and agentic capabilities. AWS is hosting Arjun…

    • Rohit Talluri Senior Worldwide Generative AI Specialist BD - Foundation Model Training & Inference
    • Arjun Desai Co-Founder, Cartesia
  • Expo Sessions Talk

    Prompt Engineering is Dead

    Thu 5 Jun, 13:15 – 13:30·Juniper: Expo Sessions

    Manual prompt crafting doesn't scale. In this session, we'll explore how to replace it with a test-driven, automated approach. You'll see how to define output evaluators, write…

    • Nir Gazit CEO @ Traceloop, OpenLLMetry co-creator
  • Expo Sessions Talk

    Optimizing inference for voice models in production

    Thu 5 Jun, 13:15 – 13:30·Willow: Expo Sessions

    How do you get time to first byte (TTFB) below 150 milliseconds for voice models -- and scale it in production? As it turns out, open-source TTS models like Orpheus have an LLM…

  • Retrieval + Search Talk

    Building Alice’s Brain: How We Built an AI Sales Rep that Learns Like a Human

    Thu 5 Jun, 14:00 – 14:20·Golden Gate Ballroom A: Retrieval + Search

    AI agents are becoming essential tools for teams of all sizes and industries - but training them to become experts in your product, business, and customerbase remains a challenge.…

  • SWE Agents Talk

    Introducing Claude Code

    Thu 5 Jun, 14:00 – 14:20·Yerba Buena Ballroom 7&8: SWE Agents

    Hear about Claude Code directly from its creator: origin story, getting your team set up, and practical tips for getting more out of Code.

    • Boris Cherny Member of Technical Staff & Creator of Claude Code
  • AI Architects Talk

    CIAM for AI: Who Are Your Agents and What Can They Do?

    Thu 5 Jun, 14:00 – 14:20·SOMA: AI Architects

    AI agents are changing the way modern SaaS products operate. Whether automating workflows, integrating with APIs, or acting on behalf of users, AI-driven assistants and autonomous…

  • Evals Keynote

    [Evals Keynote] tba

    Thu 5 Jun, 14:00 – 14:20·Golden Gate Ballroom B: Evals

    tbc

  • Security Talk

    How to defend your sites from AI bots

    Thu 5 Jun, 14:00 – 14:20·Foothill C: Security

    Constantly seeing CAPTCHAs? It used to be easy to detect the humans from the droids, but what else can we do when synthetic clients make up nearly half of all web requests.…

  • Design Engineering Talk

    AI and Human Whiteboarding Partnership

    Thu 5 Jun, 14:00 – 14:20·Foothill G 1&2: Design Engineering

    Covid sent everybody home and created the space of virtual whiteboards. At first the experience reused the physical constraints but soon it became better than a physical whiteboard…

  • Generative Media Talk

    General Intelligence is Multimodal

    Thu 5 Jun, 14:00 – 14:20·Foothill F: Generative Media

    Talking about Luma AI, our mission, and how our ML infrastructure enables SOTA multimodal model development

  • Autonomy + Robotics Talk

    Teaching Cars to Think: Language Models and Autonomous Vehicles

    Thu 5 Jun, 14:00 – 14:20·Foothill E: Autonomy + Robotics

    This session explores Waymo's latest research on the End-to-End Multimodal Model for Autonomous Driving (EMMA) and advanced sensor simulation techniques. Jyh-Jing Hwang will…

  • AI in the Fortune 500 Talk

    How to Build Agents without losing control

    Thu 5 Jun, 14:00 – 14:20·Golden Gate Ballroom C: AI in the Fortune 500

    Planning agents help solve complex tasks by breaking them into steps. They work across enterprise systems where data lives in many places. These agents are powerful but can be hard…

  • Reasoning + RL Talk

    How to Train Your Agent: Building Reliable Agents with RL

    Thu 5 Jun, 14:00 – 14:20·Yerba Buena Ballroom 2-6: Reasoning + RL

    Have you ever launched an awesome agentic demo, only to realize no amount of prompting will make it reliable enough to deploy in production? Agent reliability is a famously…

  • Reasoning + RL Talk

    A taxonomy for next-generation reasoning models

    Thu 5 Jun, 14:20 – 14:40·Yerba Buena Ballroom 2-6: Reasoning + RL

    Current AI models are extremely skilled, which was seen as the step change in evaluation scores across the industry in the first half of 2025, but often fail when presented with…

  • AI in the Fortune 500 Talk

    POC to PROD: Hard Lessons from 200+ Enterprise GenAI Deployments

    Thu 5 Jun, 14:20 – 14:40·Golden Gate Ballroom C: AI in the Fortune 500

    The transition from experimental GenAI demonstrations to robust, production-grade systems involves significant technical and organizational complexities. Humans provide a ceiling…

  • AI Architects Talk

    Testing the Un-Testable: Monitoring AI Products in the Wild

    Thu 5 Jun, 14:20 – 14:40·SOMA: AI Architects

    Evals are straightforward—like unit tests, they confirm your model got specific test cases right. But in the real world, your AI encounters millions of unpredictable…

  • SWE Agents Talk

    Software Development Agents: What Works and What Doesn't

    Thu 5 Jun, 14:20 – 14:40·Yerba Buena Ballroom 7&8: SWE Agents

    The adoption of AI into software development has been bumpy. While autocomplete tools like Copilot have gone mainstream, autonomous agents like Devin and OpenHands have generated…

  • Retrieval + Search Talk

    Building a Smarter AI Agent with Neural RAG

    Thu 5 Jun, 14:20 – 14:40·Golden Gate Ballroom A: Retrieval + Search

    RAG quality for AI agents is critical, and traditional keyword-based search engines consistently underperform in agentic or multi-step tasks, where semantic grounding and…

    • Will Bryk CEO & Co-founder, building perfect search at Exa
  • Evals Talk

    2025 is the Year of Evals! Just like 2024, and 2023, and …

    Thu 5 Jun, 14:20 – 14:40·Golden Gate Ballroom B: Evals

    AI is getting deployed without guardrails, without governance, without due diligence. Surely this is the year we’ll see a Fortune 500 CEO fired because of a preventable AI…

  • Security Talk

    How to Secure Agents using OAuth

    Thu 5 Jun, 14:20 – 14:40·Foothill C: Security

    We all know sharing passwords is bad (unless you want free TV), so why are we sharing API keys with AI? We shouldn't, and that’s why we need to talk about OAuth. In this talk,…

  • Design Engineering Talk

    tldraw computer

    Thu 5 Jun, 14:20 – 14:40·Foothill G 1&2: Design Engineering

    Learn about tldraw's latest experiments with AI on an infinite canvas. In 2024, we created tldraw computer, a loose visual programming environment where arrows and LLMs powered…

  • Generative Media Talk

    Good demos are important

    Thu 5 Jun, 14:20 – 14:40·Foothill F: Generative Media

    Creating and sharing demos is the easiest way to influence the future. It gets people to think about what's possible. A good tech demo doesn't have to be fully fleshed out. It…

  • Autonomy + Robotics Talk

    General purpose robots as professional Chefs

    Thu 5 Jun, 14:20 – 14:40·Foothill E: Autonomy + Robotics

    How we converted a bimanual robot into a professional chef that works in novel kitchens and learn new recipes from a single demonstration.

  • SWE Agents Talk

    Beyond the Prototype: Using AI to Write High-Quality Code

    Thu 5 Jun, 14:40 – 15:00·Yerba Buena Ballroom 7&8: SWE Agents

    In this case study-based keynote, Josh Albrecht, CTO of Imbue, examines the critical engineering challenges in building AI coding systems that create more than just prototypes.…

  • Reasoning + RL Talk

    Towards Verified Superintelligence

    Thu 5 Jun, 14:40 – 15:00·Yerba Buena Ballroom 2-6: Reasoning + RL

    I describe a new paradigm towards open-endedly self-improving intelligence by scaling verification to remove the human data and supervision bottleneck. The objective is to achieve…

  • Retrieval + Search Talk

    Layering every technique in RAG, one query at a time

    Thu 5 Jun, 14:40 – 15:00·Golden Gate Ballroom A: Retrieval + Search

    Start with the simplest Search - in-memory embeddings with relevance ranking. End with the most complex planet-scale Search - 70+ corpus mix of token, embeddings, and knowledge…

  • Evals Talk

    How to look at your data; what to look for, how to measure

    Thu 5 Jun, 14:40 – 15:00·Golden Gate Ballroom B: Evals

    By the end of this talk, you'll understand what it takes to apply clustering techniques and data analysis to understand what is the valuable work that your AI application is doing…

  • Security Talk

    How we hacked YC Spring 2025 batch’s AI agents

    Thu 5 Jun, 14:40 – 15:00·Foothill C: Security

    We hacked 7 of the16 publicly-accessible YC X25 AI agents. This allowed us to leak user data, execute code remotely, and take over databases. All within 30 minutes each. In this…

  • Design Engineering Talk

    Form factors for your new AI coworkers

    Thu 5 Jun, 14:40 – 15:00·Foothill G 1&2: Design Engineering

    Designing user experiences for AI means moving beyond traditional interfaces. Designers are grappling with how to create intuitive and effective interactions for these new AI…

  • Generative Media Talk

    Why you should care about AI interpretability

    Thu 5 Jun, 14:40 – 15:00·Foothill F: Generative Media

    The goal of mechanistic interpretability is to reverse engineer neural networks. Having direct, programmable access to the internal neurons of models unlocks new ways for…

  • AI Architects Talk

    The Web Browser Is All You Need

    Thu 5 Jun, 14:40 – 15:00·SOMA: AI Architects

    With the rise of MCP servers, A2A, and our trusty friend, OpenAPI, it turns out the web browser may be the default MCP server for the rest of the internet. In this talk, we'll…

  • AI in the Fortune 500 Talk

    Building Agents (the hard parts!)

    Thu 5 Jun, 14:40 – 15:00·Golden Gate Ballroom C: AI in the Fortune 500

    AI workloads are rapidly shifting from AI being used for augmentation (co-pilots), to AI becoming responsible for full, end-to-end automation (agents). But building effective…

  • SWE Agents Talk

    Ship Production Software in Minutes, Not Months

    Thu 5 Jun, 15:00 – 15:20·Yerba Buena Ballroom 7&8: SWE Agents

    Planning, coding, testing, monitoring—the endless cycle that spans 10+ tools that fragment our focus and slows delivery to a crawl. Vibe coding doesn't work when you've got 10TB of…

  • Expo Sessions Talk

    CI in the Era of AI: From Unit Tests to Stochastic Evals

    Thu 5 Jun, 15:15 – 15:30·Juniper: Expo Sessions

    Software engineers have long understood that high-quality code requires comprehensive automated testing. For decades, our industry has relied on deterministic tests with clear…

    • Nathan Sobo CEO & Co-founder of Zed, co-creator of Atom and Electron
  • Expo Sessions Talk

    Cattle, not genies: building AI agents from first principles

    Thu 5 Jun, 15:15 – 15:30·Willow: Expo Sessions

    As magical as they may seem, AI agents should be treated like any other software system. This talk will cover the best practices in designing and building AI systems including…

  • Expo Sessions Talk

    To the moon! Navigating deep context in legacy code with Augment Agent

    Thu 5 Jun, 15:15 – 15:30·Nobhill A&B: Expo Sessions

    Shortened presentation-only version of our Apollo 11 workshop

  • Expo Sessions Talk

    The Eyes Are The (Context) Window to The Soul: How Windsurf Gets to Know You

    Thu 5 Jun, 15:30 – 15:45·Nobhill A&B: Expo Sessions

    Sometimes it seems like Windsurf knows you a little too well. It's one thing to generate generic code, but to predict your next intent? From matching existing code patterns and…

  • Expo Sessions Talk

    The emerging skillset of wielding coding agents

    Thu 5 Jun, 15:30 – 15:45·Willow: Expo Sessions

    It's raining coding agents. But while many are saying they're feeling the AGI, others say they're not that useful for serious programming. How much is hype and how much is a skill…

  • Unassigned Keynote

    Trends Across the AI Frontier

    Thu 5 Jun, 16:00 – 16:20·Keynote/General Session (Yerba Buena 7&8)

    The entire AI stack is developing faster than ever - from chips to infrastructure to models. How do you sort the signal from the noise? Artificial Analysis an independent…

  • Unassigned Keynote

    Evals Closing Keynote

    Thu 5 Jun, 16:20 – 16:25·Keynote/General Session (Yerba Buena 7&8)

    The final word on Evals

  • Unassigned Keynote

    State of AI Engineering 2025

    Thu 5 Jun, 16:25 – 16:35·Keynote/General Session (Yerba Buena 7&8)

    Come hear the results of the 2025 State of AI Engineering.

  • Unassigned Keynote

    fun stories from building OpenRouter and where all this is going

    Thu 5 Jun, 16:35 – 16:55·Keynote/General Session (Yerba Buena 7&8)

    How the first LLM aggregator got started, some of the weird moments in its early growth, architecture challenges, and where we'll be taking it down the road

  • Unassigned Keynote

    Prompt Engineering is Dead - Everything is a Spec

    Thu 5 Jun, 16:55 – 17:15·Keynote/General Session (Yerba Buena 7&8)

    [!!Subject to change!!] Large models are trained through mountains of data and learned reward functions, yet - quis custodiet ipsos custodes? - what exactly are those amorphous…

ende