Details
- Starts
- Wed 1 Jul 2026, 00:00
- Ends
- Wed 1 Jul 2026, 03:30
- Type
- In person
- Category
- Tech & Startups
- City
- San Francisco
- Country
- United States
About this event
As model capabilities accelerate, evaluation is becoming the bottleneck. We can generate code at scale. We can build systems that reason through increasingly complex problems. But...
As model capabilities accelerate, evaluation is becoming the bottleneck. We can generate code at scale. We can build systems that reason through increasingly complex problems. But our ability to measure what actually works has not kept pace.
How do you know if a model genuinely improved, or if you just tuned the benchmark?
How do you eval across different coding paradigms, domains, and reasoning styles?
What happens when models reach the same outcome through fundamentally different approaches?
And how do you measure progress when there is no longer one clearly “right” answer?
Mark Hoffmann,…
Hosted by Yousra Saleh & 3 others
How do you know if a model genuinely improved, or if you just tuned the benchmark?
How do you eval across different coding paradigms, domains, and reasoning styles?
What happens when models reach the same outcome through fundamentally different approaches?
And how do you measure progress when there is no longer one clearly “right” answer?
Mark Hoffmann,…
Hosted by Yousra Saleh & 3 others
Links and contact
- Website
- https://luma.com/b33n09v9
- Registration
- https://luma.com/b33n09v9
Source
- Imported from
- Luma
- Source city
- San Francisco
- Original event
- Open source event https://luma.com/b33n09v9