Evals London

Evaluating AI applications, not just models.

A London meetup and community for people who build AI-powered products. Model benchmarks on leaderboards are interesting, but how do you evaluate your specific RAG pipeline in production?

What we tackle

🔍

RAG system quality

Evaluating retrieval quality and end-to-end RAG pipelines in production

⚠️

Production regressions

How to catch and fix problems in deployed AI systems

👥

Human-in-the-loop

Evaluation approaches that combine human expertise with automation

⚙️

Development lifecycle

Integrating evals into development and CI/CD pipelines

Who it's for

ML/AI engineers

You deploy and monitor models in production

Backend developers

You integrate LLMs or RAG into an existing product

Data scientists

You measure and improve the performance of AI systems

Product managers

You need to understand the quality of AI features

AI researchers

You study the reliability and safety of AI

Beginners welcome

Practical interest and curiosity are all you need

Meetup format

LocationDawn Capital, London
WhenWednesday 11 November 2026
Format1–3 short talks + discussion
LanguageEnglish
"What broke in production and how we fixed it"— no sales pitches

Partners

[ Venue host ]

Dawn Capital hosts edition #1.

Want to sponsor the night, or host a future edition? A room for ~100 people, a projector and some pizza go a long way.

Partner with us

Coming soon

Edition #1

Wed 11 Nov at Dawn Capital

Date and venue are set. The line-up is being put together right now, with speakers announced as each one confirms. Get on the list and you'll be the first to know when registration opens.

Wed 11 Nov 2026Dawn CapitalLondon
Get notified
A taster

What we'll dig into

  • Error analysis & failure taxonomies
  • Assertion-based evaluation
  • LLM-as-Judge frameworks
  • RAG-specific metrics
  • Evaluating agents & tool use
  • Human evaluation & annotation design
  • Wiring evals into CI/CD