← Back to Portfolio
Project 05 · Agentic AI & RAG

MeetStream
Companion

A deployed, multi-tenant AI meeting companion — it puts a voice agent into your calls, remembers every meeting in PostgreSQL + pgvector, and serves that memory back to the live agent through MCP. Each workspace's data is isolated from every other.

Overview

Most meeting bots hand you a transcript and forget everything the moment the call ends. MeetStream Companion deploys a bot that joins the call, records and transcribes it, runs the transcript through an LLM to extract structured memory — decisions, commitments, requirements, concerns — and then makes that memory queryable.

Queryable by two consumers: a live in-meeting agent that stays silent until addressed by name and can then answer "what did we decide about pricing three weeks ago?" mid-call, and a web dashboard for browsing meetings, searching memory, and managing agents.

It runs as a real product, not a local dev toy — backend and frontend are separate Railway services, and it is genuinely multi-tenant: signing up creates a brand new empty workspace or joins an existing one by join code, and every meeting, memory, document, and agent belongs to exactly one of them.

At a Glance
9
MCP tools exposed
1 date · 4 read · 4 write
384-d
Embedding dimension
all-MiniLM-L6-v2 · pgvector
10
Memory types in the schema
Decisions → unresolved questions
19
Automated tests
Unit · integration · security
Architecture

The platform reaches the companion two ways — webhooks after a call, and MCP during one. Both are authenticated per workspace.

MeetStream Platform ┌──────────────┬────────────────┬────────────────┐ │ Bots (Calls) │ MIA Agent (AI) │ Transcripts │ └──────┬───────┴───────┬────────┴───────┬────────┘ │ │ │ Webhooks MCP + chat-relay Webhooks ▼ ▼ ▼ ┌──────────────────────────────────────────────────────┐ │ FastAPI Backend · Railway service │ │ │ │ Webhook API MCP Server Meeting & Auth & │ │ (9 tools + Graph API Members │ │ chat-relay) API │ │ └──────────────┴────────────┴──────────┘ │ │ ▼ │ │ Services: MeetStream Client · Memory Extractor │ │ Embeddings · Ingestion Pipeline │ │ ▼ │ │ PostgreSQL 17 + pgvector │ │ organizations · users (members) │ │ meetings · participants · transcript_segments │ │ memories · action_items │ │ meeting_memory_embeddings (vector 384) │ │ company_knowledge_embeddings │ └──────────────────────────────────────────────────────┘ ▲ REST API (session cookie) │ React 19 + Vite Dashboard · Railway Day view · Notebook · Agent · Members · Auth
Key Features
Debugging Note — Agent Silence

The bot would join, record and transcribe perfectly, but the agent itself never produced output — across three different model providers. That pattern pointed hard at a platform bug.

It wasn't. create_bot() was manually injecting socket_connection_url and live_audio_required into the bot-creation payload whenever an agent config was set. Those look like internal bridge endpoints that MeetStream wires up server-side; overriding them broke the agent's audio/text bridge silently. The tell was that the same agent config deployed fine through MeetStream's own API Playground — so the difference had to be in our request, not their platform. Removing both fields fixed it.

Security Note — Closing the Tenant Boundary

The first version had one shared workspace, which was fine while I was the only user. Opening signup turned that into a real data-exposure problem overnight: anyone who created an account could see every meeting already in the system. I closed signup the same day and rebuilt the data model around genuine tenancy — every meeting, memory, document and agent scoped to an organization_id resolved from the caller's session.

The subtler bug surfaced afterwards. Agent activation, agent updates, and the chat relay all accepted a client-supplied agent or meeting ID and performed no ownership check — so one workspace could hijack or overwrite another workspace's live agent, or post into another workspace's meeting chat. Listing agents had the mirror-image flaw, showing every agent on the underlying MeetStream account rather than the ones this workspace created. Both are now ownership-verified server-side, and there is an integration test suite (test_tenant_isolation_integration.py, test_security.py) whose whole job is to fail if that boundary regresses.

Retrieval Note — Recall the Index Was Quietly Dropping

Search kept missing results a human could see were relevant — not ranking them low, omitting them. The ranking layer looked fine, so the fault was upstream of it.

The vector index is ivfflat with lists=100, and pgvector defaults to probes=1 — each query scans roughly 1% of the index's clusters. That's a sane trade at the 100k-row scale the index is tuned for; at a single workspace's real data volume it means genuinely relevant rows are never examined at all. Raising probes to 10 buys back the recall for a small, bounded amount of query time, and still scales as the table grows.

The second half was semantic rather than mechanical. Asking "what did we decide on September 2" was surfacing a meeting from the 3rd where someone recapped the 2nd, ahead of the actual September 2 meeting — both are textually about that date, so pure similarity can't separate them. Results now get a stable date-match boost, and relative words like "yesterday" are resolved when content is indexed, against the meeting's own date rather than the moment of the search. Resolving them at search time instead would make every old transcript saying "yesterday" collide with any query that happens to say it too.

Tech Stack
Backend
Python 3.12 FastAPI Uvicorn SQLAlchemy (async) Pydantic
Data & Retrieval
PostgreSQL 17 pgvector asyncpg sentence-transformers all-MiniLM-L6-v2 Reciprocal Rank Fusion
Agent & LLM
MCP (JSON-RPC 2.0) MeetStream MIA Groq OpenAI Anthropic
Auth & Security
Session cookies Per-workspace bearer tokens HMAC-signed webhooks Replay protection Org-scoped repositories
Frontend
React 19 Vite 8 oxlint
Ops
Railway Docker Cloudflare Tunnel (dev) pytest
Interface