Podium Coach
back to projectsImprove your live presentation with a real-time coaching app.
Overview
Podium Coach is a real-time coach for public speakers. It listens to the presenter, observes the room, then offers simple guidance to help better connect with the audience.
The Problem
So many public presentations are given every day, and many are weaker than they should be. Often the problem is not the material, it's the actual real-time presentation. Once a talk is under way, it can be hard for another person to coach the presenter. For instance, the presenter may not hear their own filler words, or notice the clock, or keep their eyes on the audience to gauge interest.
The Solution
Podium Coach is a camera-enabled app that can listen to the speaker while also watching the audience. The presenter's audio is measured for speaking pace, volume, filler words, and time remaining against the plan.
Meanwhile, a camera faces the audience and samples the room. A vision model measures how many people are looking towards the speaker, and reports a positive or negative attention-trend to the speaker with a simple graphic.
Both signals meet in the 'brain' of the app which analyzes the input and offers simple tips to the presenter such as: Ask a question. Speak louder. Land it — you're over time.
Design Concepts
These simple coaching cues help remove the self-monitoring pressure from the presenter so they can relax more and get into the flow to connect with their audience.
Code decides when to speak, from fixed numeric thresholds. The model only decides how. If a model call fails, a fallback line for that category is shown instead, so the screen at the podium never displays an error.
A concept by Adam Burgh, built together with Reed O'Beirne.
See it for yourself at GitHub.
Context
Podium Coach was built in four hours at the AI Tinkerers Agents, Everywhere global hackathon in Seattle, September 2026. It reached a working prototype inside that window and remains at that stage; enough to demonstrate the idea end to end, not a finished product.
The concept came from Adam Burgh, who has run many hackathons and group presentation sessions and has watched hundreds of people give talks — and seen how little anyone at the back of a room can do to help in the moment.
Technical
| Listening | Presenter audio captured in ten-second chunks and transcribed by Deepgram nova-3 with word timestamps and filler-word detection. A Whisper model over OpenRouter stands by as transcription failover. |
|---|---|
| Watching | A phone facing the audience captures one frame every so often. Claude Haiku 4.5 rates each frame against a fixed attention rubric and returns a score, never an identity. The image data is not stored. |
| Deciding | Deterministic thresholds choose when a cue fires and which category it belongs to: words per minute, fillers per minute, seconds since the last pause, section time against the outline, and change in engagement. The thresholds live in a configuration file, not in code. |
| Writing | Claude Sonnet 5 writes the chosen cue in six words or fewer. A fixed fallback line exists for every category, so a failed model call degrades to a plain sentence rather than an error. |
| Output | A phone on the podium polls for state every two seconds and shows one cue, a countdown, and a single arrow for the direction of audience attention. After the talk, a written recap of the run. |
| Built with | Python, Deepgram nova-3, Anthropic Claude Haiku 4.5 and Sonnet 5, OpenRouter, Cloudflare Tunnel, Ambiguous.ai, OpenAI |
Privacy
Pointing a camera at an audience deserves a clear answer, so here it is. Audience frames are sampled regularly, sent to a vision model with a fixed rubric that rates posture and attention, and discarded immediately. No frames are stored. No faces are detected and no one is identified.
The presenter's own audio is transcribed and kept only for the post-talk recap that the presenter receives. No data is preserved outside the app.