AI Interviewer: Gemini Live & LiveKit
A Proof-of-Concept demonstrating AI-powered preliminary job interviews in real-time using Google's Gemini Live API and Livekit for natural, conversational interactions.

Overview
This proof-of-concept shows how AI agents can conduct preliminary job interviews in real time—natural, conversational, and tailored to a specific role. A recruiter enters the candidate's name and a job description through a simple Next.js interface, which spins up a Python-based Livekit agent powered by Google's Gemini Live API. The agent greets the candidate, asks role-relevant questions one at a time, follows up on answers, and gracefully wraps up within a preset time limit—all with low-latency, back-and-forth voice. As AI prototyper and developer, Rohan integrated the Gemini Live realtime model, set up Livekit's real-time audio infrastructure, built the Python interview agent, and developed the Next.js frontend that launches each session.
The problem
Traditional preliminary interviews can be time-consuming and resource-intensive. This project aims to automate and streamline this process using real-time AI agents.
Key features
- AI-Driven Conversational Interviews (Gemini Live)
- Real-time Audio/Video Infrastructure (Livekit)
- Structured Interviewing based on Job Descriptions
- Time-Bound Conversations
- Simple Next.js User Interface for inputting candidate/job details
What I did
- Integrated Google's RealTime Model Gemini Live API for AI voice and intelligence, set up Livekit for real-time communication, developed the Python-based Livekit Agent, and built the Next.js frontend to initiate interviews
AI under the hood
The interviewer is driven by Google's Gemini Live API, a realtime multimodal model that processes and generates speech with low enough latency for fluid, interruptible conversation. The agent is configured with dynamic instructions—the candidate's name, the job description, a question strategy, and a strict time budget—so a single LLM both understands spoken answers and decides relevant follow-ups on the fly. Livekit's agent framework supplies the real-time audio transport and manages the AI as a participant in the call, while a Next.js frontend injects the per-interview context. It's a clear example of conversational, instruction-steered AI operating in a live, time-bound setting.