← FOG·CITY

“ Talks

Bay Area Frontier Research Club #25 | Recursive Self-Improvement (dinner + paper discussion @ The Residency)

When
Wednesday, September 30 · 5:30 PM – 8:30 PM
Listed by
STDISF — Luma
Luma discover page Event Link Recursive self-improvement as an empirical research program. The claim under examination is narrow: systems that improve the process by which they themselves improve — not merely agents that retry within a task, and not merely human-supervised iteration at higher throughput. That distinction matters, because most of what is currently labeled RSI sits somewhere between those two. This session treats the problem as a stack with separable research questions: the models and methods that generate improvement, the feedback and evaluation that verify the improvement is real, and the infrastructure that lets the loop run without a human in the middle. Talks stay short. Discussion is the point. Papers and materials are circulated in advance. Speakers are researchers operating inside these loops; the room is expected to press on assumptions, measurement, and what would count as decisive evidence. Co-hosted with Inventors Residency of San Francisco. ―――――――――― The Frontier Research Club is a curated forum for rigorous, technical discussion at the frontier of AI. We convene researchers from the frontier labs, Stanford, Berkeley, and the teams building in production to examine concrete work — papers, methods, and results — with a bias toward assumptions, evaluation methodology, failure modes and convincing evidence. Presentations are intentionally brief so the majority of time is reserved for questions and critique. Materials are shared in advance so the conversation starts at depth. Agenda 5:30pm: Doors open 5:30pm – 6:30pm: Networking + light dinner 6:30pm – 8:00pm: Research presentations + discussion 8:00pm – 8:30pm: Networking ―――――――――― Presenters & topics Talk 1: FutureSim — Replaying World Events to Evaluate Adaptive Agents Shashwat Goel — PhD Researcher, ELLIS / Max Planck Institute Tübingen · Best Paper, ICML Forecasting Workshop · AAAI Outstanding Paper Award Shashwat works on the science of evaluations and how AI can iteratively improve — the exact hinge of this session. His PhD at Max Planck, advised by Jonas Geiping and co-mentored by Douwe Kiela (CEO, Contextual AI), focuses on designing evaluations and methods to scale AI supervision. Before Tübingen he contributed to Representation Engineering and to WMDP, now the standard benchmark for unlearning dual-use knowledge, and his earlier work earned an AAAI Outstanding Paper Award — top 3 of more than 12,000 submissions. If an agent is improving itself, how would we know — and can today's frontier agents actually update their beliefs as the world changes? FutureSim tests that question directly. The environment replays real-world events in chronological order past the model's knowledge cutoff: agents receive daily news over a three-month simulation, maintain a portfolio of forecasts, and decide entirely on their own when to update which beliefs — long-horizon and open-ended, yet reproducible and grounded in real event data. The results separate frontier agents starkly: the best reaches 25% accuracy, several score worse than making no prediction at all, and models with stronger priors start ahead but barely improve as evidence accumulates — scale buys better starting knowledge, not better updating. Shashwat will present the benchmark's design, what the gaps reveal, and the research it makes answerable: test-time adaptation, epistemic humility, memory, search, inference-scaling, and multi-agent self-play. Pre-read: Goel et al., FutureSim: Replaying World Events to Evaluate Adaptive Agents, Best Paper, ICML Forecasting Workshop. ―――――――――― Talk 2: TBA ―――――――――― Lightning talks — The Residency RSI cohort Three short talks (3 minutes each) on work in progress from researchers living and building at The Inventors Residency. ―――――――――― Want to present your work? If you have a research paper you’d like to discuss at one of our next sessions, please submit it for consideration. Submit your paper here ―――――――――― Who should attend • Researchers working on self…

More talks soon