- AI Engineering Keynote
Serving 70B models on a budget
Wed 12 May, 09:30 – 10:15·Main Stage
A working account of cutting inference cost by an order of magnitude without giving up latency: continuous batching, speculative decoding, and the three quantisation choices that actually mattered. Includes the two approaches that lost us a month.
Priya Raman Principal Engineer, Northwind Labs
- Platform & Infra Talk
Four hundred engineers, one repository
Wed 12 May, 09:30 – 10:00·Room 2B
What actually breaks at that size, in the order it breaks: code ownership, CI queueing, and the review culture nobody wrote down. Two things we would do again and one we would not.
Dmitri Sokolov Infrastructure Architect, Ravenline
- Platform & Infra Talk
Feature flags are a database problem
Wed 12 May, 11:00 – 11:30·Main Stage
Every flag system starts as a boolean and ends as a query planner. How we cut evaluation latency to microseconds, and why the interesting part turned out to be deletion, not rollout.
Nadia Farouk Senior Software Engineer, Ostmark
- Amara Nwosu Product Engineer, Ostmark
- Platform & Infra Talk
Your build is slow because of four things
Wed 12 May, 11:00 – 11:30·Room 2A
Build times decay for boringly consistent reasons. We instrumented ours for a year; this is what we found, in order of how much time each cost, and what fixing them actually took.
Marcus Okafor Staff Platform Engineer, Meridian Systems
- AI Engineering Talk
Prompt injection in production: a field report
Wed 12 May, 14:00 – 14:30·Room 2A
Six months of attempts against a customer-facing assistant, what got through, and the four mitigations that survived contact. No threat-model diagrams — just the payloads and what they cost us.
Tomiwa Adeyemi Security Engineer, Northwind Labs