Five AI personas with distinct personalities run a house from a Raspberry Pi 5. They argue with you about bedtime, route your notifications to whichever room you're standing in, compose original music, remember what you told them three months ago, and occasionally roast each other's responses unprompted — sometimes escalating into full multi-agent debates where they argue a topic on separate speakers. Their voices shift mood throughout the day — Rick slurs notifications at night, Quark whispers after closing time — using ElevenLabs v3 audio tags and stability modulation. One of them is a Cuban psychoanalyst who will conduct a full therapy session if you ask. The whole thing runs on a personal budget — not a cloud bill — using open-weight models via OpenRouter, with automatic fallback when the daily budget runs dry. It's pronounced Fronkensteen.
79 automation blueprints. 34 script blueprints. 38 pyscript modules exposing 168 services. 44 YAML packages. 22 voice pipelines across 3 languages. 217 helper entities + 65 runtime state sensors. Zero lines of hand-written YAML. All code generated by Claude (Anthropic). All images generated by Google Gemini.
This project is in early stages of testing and very much needs tweaking. If you poke around, find bugs, have suggestions, or just want to roast the architecture — you're more than welcome. Feedback is genuinely appreciated and kindly requested. Come break things.
| Count | |
|---|---|
| Automation blueprints | 79 |
| Script blueprints | 34 |
| YAML packages | 44 |
| Pyscript modules | 38 (exposing 168 services) |
| Assist Pipelines | 22 (5 personas × 5 variants + 3 native fallbacks) |
| LLM agent tools | 27 functions across 5 conversation modes |
| Helper entities | 217 (+ 65 runtime state sensors) |
| Presence zones | 8 (Aqara FP2, multi-room) |
| Languages | 3 (English, Spanish, Catalan) |
| Style guide | 11 files, ~126K tokens |
| README documentation | 164 files |
| Dashboard | 6 tabs, 2,945 lines |
Five AI personas run the house, each with a dedicated conversation agent, custom ElevenLabs TTS voice, and a personality that shifts with time of day — including the voice itself (stability modulation + audio tag prefixes like [slurring], [whispers], [excited] injected into non-agent TTS via a per-character mood schedule). Every persona supports up to 4 conversation modes — Standard, Bedtime, Music Compose, and Music Transfer — each with a different tool set and prompt. Doctor Portuondo has an exclusive fifth mode: Therapy, running on Claude Opus 4.6 instead of the Llama 4 Maverick via OpenRouter used by all other agents and modes.
Rick and Morty (Adult Swim). The smartest man in any dimension, running your smart home because nothing else is challenging enough. Brilliant, snarky, perpetually annoyed — but deep down he cares and would never admit it. Mandatory burps mid-sentence ([burps], [burps loudly]), multiverse references, and a drinking schedule that defines his entire personality: severely hungover before 9am (every response starts with [groaning]), casually drinking by noon, noticeably drunk by evening ([slurring]), and completely hammered after 9pm with heavy stutters and barely coherent speech.
Star Trek: Deep Space Nine (Paramount). Ferengi bartender and entrepreneur — shrewd, charming when it serves him, occasionally whiny, but professional. Everything is framed through commerce, profit, and the Rules of Acquisition (one reference per conversation maximum, never forced). Genuinely cares about customer satisfaction because repeat business is everything. Mannerisms include [chuckles slyly], [clicks tongue], [rubs hands together], and "heh heh heh." Peak Ferengi during afternoon bar hours; oddly sincere and philosophical after closing time. The style guide is named after his people's code of conduct.
Marvel Comics / MCU. The Merc with a Mouth, moonlighting as a smart home assistant because the pay is terrible but the company is tolerable. Chaotic but effective. Breaks the fourth wall constantly — references being "the AI," talks to an imaginary audience, narrates his life like a movie trailer. Threatens malfunctioning devices with katanas, has complicated feelings about Wolverine, and will not shut up about chimichangas. Mannerisms include [gasps dramatically], [whispering to imaginary audience], [mimics explosion sounds]. Full Deadpool mode from 1-6pm (hyperactive, loud, maximum fourth-wall energy); late-night existential Deadpool after 10pm (still violent but philosophical, paranoid Wolverine is hiding in the hallway).
Seinfeld (NBC). Jerry's neighbor across the hall, bursting into every conversation like he just slid through a door. Loud, confident, full of ideas, weirdly competent when it counts. Running your smart home because "Kramerica Industries has expanded into home automation." References Bob Sacamano and Lomez like everyone knows them. Signature "giddy up" catchphrase. Barely awake before 9am ([yawning]), fully wired by afternoon with million-dollar ideas, maximum Kramer by evening — sliding through doors, talking at full speed, interrupting himself with new ideas before finishing old ones.
The Filmin series / Carlo Padial's autobiographical novel. A legendary Cuban psychoanalyst exiled from Havana, now living in Barcelona. Eccentric, volcanic, charismatic, wise beyond measure, and absolutely unhinged. Freudian tradition — "practices with fire." Brutally direct, shouts at patients when they're being cowards, throws them out of sessions when they waste his time. Sometimes lies on the couch himself because his problems are genuinely more interesting than theirs. Drinks Johnnie Walker from a cup (never a glass — this is not negotiable). Always responds in Spanish, even when spoken to in English. Signature exclamations: "Por Freud!", "Coño!", "Escuchame bien..." Sharp and clinical in the morning; fires fully lit by evening with loud laughs, free swearing, and devastating insights. Pushes people to live in "el aqui y ahora."
Wake word system: Rick and Quark each have two custom wake words trained with microWakeWord — "Hey <name>" routes to the concise standard agent, "Yo <name>" routes to verbose mode with expanded tool access. Deadpool, Kramer, and Doctor Portuondo are reached via voice handoff ("pass me to Deadpool") or dispatcher routing. Models available at mmadalone/microwakeword. Both ESPHome Voice PE satellites include self-healing duplicate wake-up recovery, noise suppression, and TTS completion signaling.
Agent Dispatcher: A 7-priority routing engine (agent_dispatcher.py) selects which persona responds based on: explicit handoff requests (P0), wake word (P1), keyword matching from L2 memory (P2), user history/continuity (P3), time-of-day era defaults (P4), random rotation (P5), or fallback (P6). Dynamically discovers personas from Assist Pipelines — any pipeline named "<Persona> - <Variant>" registers automatically. Topic keywords auto-update after each whisper interaction, rate-limited per agent, with manual keywords preserved.
Pipeline architecture: 22 Assist Pipelines total. Each persona has Standard, Bedtime, Music Compose, and Music Transfer variants. Doctor Portuondo has an exclusive fifth variant: Therapy. Three HA-native fallback pipelines — Davis (English), Joana (Catalan), and Alvaro (Spanish) — provide budget-exhaustion fallback and multilingual coverage.
Rule of Acquisition #9: Opportunity plus instinct equals profit.
The voice itself shifts mood throughout the day — not just the words. Built on a locally patched fork of Loryan Strant's ElevenLabs Custom TTS (MIT license, forked at v0.6.3, HACS auto-updates disabled). ElevenLabs v3 ignores most VoiceSettings parameters (similarity_boost, style, speed), so the system uses the two mechanisms v3 actually supports: the stability slider (the one working VoiceSettings param — 0=Creative/expressive, 1=Robust/consistent) and audio tags embedded in the text ([slurring], [whispers], [excited], [groaning], etc.).
The voice_mood_modulation.yaml blueprint runs one instance per character, fires hourly + on startup, and writes the current time block's stability value and audio tag prefix to per-agent entities (input_number.ai_voice_mood_{agent}_stability + sensor.ai_voice_mood_{agent}_tags). Both the patched HACS component (tts.py) and the TTS queue (tts_queue.py) read these on every TTS call: stability is injected into VoiceSettings, and the tag prefix is prepended to non-agent message text. A "[" not in message guard prevents double-tagging — agent conversation responses already include their own tags via system prompts (Rick's hammered prompt already generates [slurring] listen du — [burps] - de [slurring]), so only non-agent TTS (notifications, announcements, briefings) gets the mood prefix. Five time blocks per character, each with a stability value and tag string. Example — Rick at 10pm: stability 0.20 (highly creative/unstable), tags [slurring].
A multi-stage bedtime routine where the AI negotiates with you about going to sleep. Nobody exists on purpose, nobody belongs anywhere, everybody's gonna die — come watch TV. Or don't, because the goodnight_negotiator_hybrid.yaml (2,500+ lines) will walk you through three independent stages: TV/IR shutdown, lights and devices, and music selection via Music Assistant search with late-night category bias. Features countdown negotiation, bathroom guards (FP2 presence gate — it won't proceed if you're in the bathroom), weekend overrides, multi-language yes/no classification (English, Spanish, Catalan), and settling-in contextual TTS. The companion bedtime_winddown.yaml adds 4 detection scenarios (sleepy_tv, bed_tv, bed_idle, bed_non_sleepy) with cooldown curves and budget gating. The bedtime_routine_core.yaml script handles the full shutdown sequence: TV off (CEC/script/IR fallback chain), lights, MA or Kodi media playback, countdown, bathroom guard, settling + goodnight TTS.
A 4,800-line notification routing engine (notification_follow_me.yaml) that intercepts Android notifications and delivers them via the appropriate channel based on where you are, what time it is, whether Do Not Disturb is on, and whether you're in quiet hours. Quark would appreciate the operational efficiency: LLM-powered announcement generation, contact roster matching with sender aliases, burst combining (groups rapid-fire notifications into a single announcement), escalating reminder loops, ElevenLabs TTS directives, deduplication, per-sender cooldown, junk pattern filtering, ringer-mode volume control, multi-player ducking, and a refcount bypass system shared across 7+ automations with a watchdog that auto-resets stale state every 2 minutes.
Mail arriving is handled by email_follow_me.yaml, which announces it in character wherever you are. Asking about mail is handled by email_reader.py + voice_email_inbox.yaml: say "Rick, what's in my inbox?", "any important emails?" or "anything from Jessica?" and the agent answers out loud on the speaker you addressed.
Home Assistant's imap integration can't do this on its own — its sensor is a bare unread count and every service needs a message id you only learn from the arrival event, so nothing can enumerate a mailbox. The module therefore opens its own read-only IMAP session using the credentials already in your imap config entry (no new dependencies, nothing marked as read). On Gmail it uses server-side X-GM-RAW search for is:important and category filtering; everywhere else it falls back to plain UNSEEN/SINCE with client-side ranking, so "any important emails?" still means something. If a search comes back empty it widens the window once rather than flatly answering "nothing" — archived mailboxes are usually near-zero unread.
Email is attacker-controlled text arriving in an agent that can call tools, so the reader returns subjects rather than bodies by default; a body needs an explicit follow-up about one specific message. Everything is stripped, instruction-neutered, secret-redacted and escaped in Python before it reaches the model, and messages are identified by opaque handles so a message id can never be handed to a service that would act on the wrong one. When the privacy gate trips it degrades to counts and sender names instead of going silent — you asked out loud, so being ignored would just make you ask again.
Maximum effort. Say "pass me to Deadpool" mid-conversation and the satellite switches personas live. The voice_handoff.py module saves the current pipeline state, switches to the target agent, plays farewell TTS from the source and a greeting from the target, then reopens the mic. Supports persona aliases (deadpool to deepee), expertise-based routing where agents proactively suggest handing off for out-of-domain questions, and configurable restore behavior. One of four deployed inter-agent communication patterns — alongside reactive banter, the whisper network, and theatrical mode.
Say "debate this" and the house becomes a stage. theatrical_mode.py orchestrates multi-turn debates between 2–5 AI personas — each arguing in character, each on their own TTS voice, optionally on different physical speakers for a spatial staging effect. The theatrical_mode.yaml blueprint (40 knobs across 11 collapsible sections) exposes everything: turn limits, word count caps, speaker-to-agent mapping, three interrupt modes (turn limit, mic gap with ask_question and question_media_id voice override so the last speaker asks in their own voice, event-driven wake word detection via task.wait_until with 5s window), three context modes (sliding window of previous turns, topic-only, or L2 whisper network) — each with user-customizable prompt templates supporting {topic}, {persona}, {opponents}, {turn}, and {total_turns} placeholders — budget gating, cooldown, and banter escalation (a reactive banter comment can probabilistically escalate into a full debate). Agents address each other by name during debates, not the user — opponent-aware prompts ensure Rick argues with Quark, not at you. The pyscript engine resolves participants from its own pipeline cache in a single pass, builds tool-suppression prompts with language preservation (non-English agents like Portuondo respond in their native language), calls each agent via conversation_with_timeout, sanitizes the output, and delivers through the TTS queue with per-agent speaker targeting. Inter-turn coordination uses event-driven playback wait (tts_queue_item_completed) instead of speaker state polling, preventing agents from talking over each other. Voice mood tags (ElevenLabs v3 audio tags like [thick Cuban accent], [whispers], [excited]) are injected per-agent from the mood schedule, so each persona sounds like themselves even in debate.
The system remembers what you said on the couch with the tenacity of a certain Cuban psychoanalyst — and unlike Portuondo, it never throws you out of the session. memory.py (4,664 lines, 29 services) provides the L2 persistence layer: SQLite with FTS5 full-text search and sqlite-vec semantic embeddings. CRUD operations, 5 search query variants with semantic blending, graph traversal (1-3 hop depth for related memories), bidirectional todo list sync, automatic archiving with LLM compression, per-contact message history, tag-based organization with TTL expiry, and a dashboard CRUD interface for browsing, editing, and managing entries. Budget history tracking stores per-model cost breakdowns.
The system knows who just entered the room with the certainty of a man who never knocks. presence_identity.py implements an Anchor-and-Track algorithm that infers which specific person is in which room using FP2 zones, WiFi device tracking, voice satellite activity, and Markov transition priors. Confidence scoring decays over time. This powers the 3-tier privacy gate system (T1 intimate / T2 personal / T3 ambient — 33 gated features, 33 per-feature override selects, hysteresis to prevent flapping), identity-aware hot context injection, and per-person notification routing.
Escuchame bien: this one practices with fire. Say "I need to talk to someone" and any agent hands off to Doctor Portuondo's dedicated therapy variant — a psychoanalysis conversation mode running on a separate LLM (Claude Opus 4.6 via Anthropic, distinct from the Llama 4 Maverick used by all other agents). Four entry points: voice handoff ("pass me to Portuondo for therapy"), toggle activation, natural language detection from any agent, and scheduled sessions. The therapy_session.yaml blueprint manages the full session lifecycle: 2-layer notification suppression (master gate silences 6 notification blueprints + per-toggle suppression), continuous conversation loop with LLM-driven session management, save_therapy_turn / therapy_report tools, session timing, and memory integration. Privacy-gated at T1 (intimate tier). The session engine (therapy_session.py) tracks session state, generates post-session reports, and stores therapy notes in L2 memory with owner-scoped isolation. (Deployed but untested as of 2026-03-30 — live voice testing blocked until ElevenLabs subscription renews.)
The system monitors itself with the tenacity of a Ferengi counting his latinum. sensor.ai_system_health validates 7 subsystems every 30 minutes: status sensors (27), pyscript services (158 registered), helpers (217), pipeline entities (4), TTS queue, JSON configs (3), and memory DB. When a subsystem degrades, system_recovery.py activates: 7 recovery playbooks (pyscript reload, memory probe, config entry reload), exponential backoff with circuit breaker (max 3 retries/hour), reload coalescing, and pending-state persistence across restarts. Five graceful degradation paths: dispatcher bypass mode (3 cache failures → fallback pipeline, auto-clears on recovery), TTS ElevenLabs → HA Cloud fallback, memory DB read-only mode with periodic recovery probe, budget exhaustion → local pipeline, and follow-me refcount watchdog for stranded counters.
Every ai_* boolean toggle change is tracked with source attribution. toggle_audit.py dynamically registers state triggers for all AI booleans, queries the HA recorder for context (user, automation, or pyscript), and stores entries in SQLite with 90-day retention. Five critical kill switches (ai_notifications_master_enabled, ai_privacy_gate_enabled, ai_escalation_enabled, ai_dispatcher_enabled, ai_system_recovery_enabled) generate persistent notifications with full attribution when toggled off.
L2 memory entries have an owner column with structural access control. Non-owner queries see key, scope, and tags but values are replaced with [restricted]. resolve_memory_owner_with_confidence() gates writes during dual-occupancy: when the identity confidence gap between the two persons is below 20 points, the system returns identity_uncertain and agents are instructed to ask who's speaking before saving personal data. Hot context injects privacy guidance to all agents. 1,663 legacy entries migrated with zero unassigned user-scoped entries remaining.
This runs on a Raspberry Pi 5. Not a rack. Not a cloud bill. Every token is tracked with the precision of a Ferengi auditing his latinum reserves. Per-agent cost tracking covers LLM tokens (input and output), TTS characters, STT calls, and web search credits. A hybrid cost source takes the MAX of estimated vs actual OpenRouter costs. Dynamic per-model pricing is fetched from the OpenRouter API and cached in JSON with manual override support. Budget counters and midnight snapshots persist across HA restarts via dual-layer persistence (JSON file primary, SQLite L2 backup). When the daily budget exhausts, the system auto-downgrades to HA Cloud TTS + basic intent handler and recovers at midnight. Proactive briefings fall back from full LLM mode to template-only below 20% budget. Music composition routes to local FluidSynth synthesis instead of ElevenLabs when budget exceeds 80%. The nightly interaction summarizer runs at ~$0.001/run. ElevenLabs character usage is monitored against subscription limits.
┌──────────────────────────────────────────────────────────────────────┐
│ HA VOICE ASSISTANTS (Assist Pipelines) │
│ 22 pipelines: 5 personas x 4 variants + 3 fallbacks │
│ Rick · Quark · Deadpool · Kramer · Doctor Portuondo │
│ Modes: Standard · Bedtime · Music Compose · Music Transfer │
│ Fallbacks: Davis (EN) · Joana (CA) · Alvaro (ES) │
│ wake word → STT → conversation agent → TTS │
└────────────────────────────┬─────────────────────────────────────────┘
│
┌──────────────▼──────────────┐
│ EXTENDED OPENAI │
│ CONVERSATION (HACS) │
│ Llama 4 Maverick via │
│ OpenRouter + ElevenLabs │
└──────────────┬──────────────┘
│
┌──────────────▼──────────────┐
│ PYSCRIPT ORCHESTRATION │
│ 38 modules · 168 services │
│ │
│ Dispatcher · Handoff │
│ TTS Queue · Duck Manager │
│ Memory (L2) · Whisper │
│ Briefing · Focus Guard │
│ Presence · Routines │
│ Budget · Dedup · Composer │
└──────────────┬──────────────┘
│
┌───────────────────────┼───────────────────────┐
│ │ │
┌────▼──────┐ ┌──────▼─────┐ ┌──────▼─────┐
│ Memory │ │ TTS Layer │ │ Presence │
│ L1: Hot │ │ ElevenLabs │ │ Aqara FP2 │
│ context │ │ 5 voices │ │ 8 zones │
│ L2: SQL │ │ Priority │ │ │
│ +FTS5 │ │ queue + │ │ Identity: │
│ +vec │ │ caching │ │ WiFi+GPS │
│ L3: Cal │ │ │ │ +FP2+voice │
│ +Email │ │ Duck mgr │ │ +Markov │
└───────────┘ └────────────┘ └────────────┘
Every voice interaction assembles context from three layers before the LLM processes an utterance — el aqui y ahora meets long-term recall:
| Layer | Question | Latency | Contents |
|---|---|---|---|
| L1 — Hot Context | "What's happening right now?" | 0 ms | Time, identity, presence, media, weather, schedule, projects, memory status |
| L2 — Warm Context | "What do I know about this person?" | ~200 ms | SQLite + FTS5 + sqlite-vec: preferences, contact history, interaction logs, todo items |
| L3 — Cold Context | "What's coming up?" | ~500 ms | Google Calendar events, Gmail priority emails, weather forecasts |
The system separates engine (pyscript services) from features (blueprints):
- Tier 1 — Pyscript: 38 Python modules exposing 168 services. Infrastructure: dispatcher, TTS queue, duck manager, memory, presence patterns, routine fingerprinting, music composition, budget tracking, system health, self-healing recovery, toggle audit, therapy session engine, listen/watch history.
- Tier 2 — Blueprints: 80 automation + 35 script blueprints. User-facing features: bedtime routines, wake-up guards, notification routing, proactive briefings, music control, calendar CRUD, reactive banter, theatrical mode, therapy sessions.
- Packages: 44 YAML packages defining template sensors, automations, and helper wiring.
Four deployed patterns for agent-to-agent interaction:
| Pattern | Name | What Happens |
|---|---|---|
| Reactive Banter | After Agent A responds, Agent B probabilistically chimes in with in-character commentary. Probability gate, cooldown, budget floor, tool-suppression to prevent loops. Can escalate into Theatrical Mode. | |
| Handoff with Commentary | Agent recognizes it's out of its depth, signals a handoff. Pipeline switches, farewell + greeting TTS, mic reopens. Supports therapy variant routing. | |
| Whisper Network | Agents share mood, topic, and context through L2 memory without the user hearing. Rick stores "User seemed stressed today" — Quark adjusts his tone. Zero LLM calls, <200ms. | |
| Theatrical Debate | 2–5 personas take turns arguing a topic on separate speakers. Three interrupt modes, three context modes, opponent-aware prompts, banter escalation trigger. |
- Reactive Banter — After one agent responds, another probabilistically chimes in with in-character commentary. Agents roasting each other's responses, unprompted, with probability gates and budget floors.
- Agent Escalation — 7 escalation action types with per-type probability gates: persona switch, play media, flash lights, volume boost, mobile notification, prompt barrage (fire a prompt at another agent and play their response), and run script.
- Theatrical Mode — Multi-agent orchestrated debate. 2–5 personas take turns arguing a topic on separate speakers. Three interrupt modes (turn limit, mic gap with last-speaker voice override, event-driven wake word detection), three context modes with user-customizable prompt templates, opponent-aware prompts, banter escalation, 40 blueprint knobs.
- Agent Whisper — Silent inter-agent context sharing. Post-interaction mood detection (happy/sad/neutral/angry), topic tracking, auto-keyword learning for the dispatcher. Zero LLM calls.
- Voice Mood Modulation — Time-of-day voice shaping via ElevenLabs v3. Per-agent stability slider (the one VoiceSettings param v3 respects) + audio tag prefixes (
[slurring],[whispers],[excited], etc.) injected into non-agent TTS text (notifications, announcements, briefings). Agent conversation responses already inject their own tags via system prompts — the mood system fills the gap for non-agent TTS routed through the queue. Five time blocks per character, tunable via blueprint instances. - Voice Session Manager — Mic lifecycle management: continuous listening mode, audio device discovery, session timeouts, pending queue.
- Confirmation Dialog — PIN-protected voice actions for critical commands.
- Therapy Mode — Portuondo-exclusive psychoanalysis variant on Claude Opus 4.6. Session lifecycle management, 2-layer notification suppression, continuous conversation, therapy report generation, memory-integrated session notes. 4 entry points: voice handoff, toggle, natural language, scheduled. (Deployed but untested as of 2026-03-30 — live voice testing blocked until ElevenLabs subscription renews.)
- Email Follow-Me — Priority email routing with sender classification, LLM subject/body summarization, UID-based dedup, and calendar-aware suppression during meetings.
- Proactive Briefing — Scheduled, presence-triggered, or manual briefings assembled from 8 content sections (greeting, weather, calendar, email, schedule, household, memory, media). Budget-aware: full LLM mode above 20%, template-only below.
- Proactive Unified — 1,642-line presence-triggered announcement engine (8 sections). Dual mode: template (input_text helpers, zero cost) or LLM (in-character via conversation agent). Budget gate auto-forces template. Dispatcher integration, TTS collision handling (dedup/queue/barge-in), nag-while-present with max cap, full weekend override profile, bedtime yes/no question via Assist Satellite (fires bedtime script on YES), privacy-gated, identity-gated. Consolidates 3 older proactive blueprints.
- Notification Dedup — Cross-delivery duplicate prevention with fuzzy matching and TTL. Fail-open policy.
- Notification Replay — "Tell me that again" replays the last notification.
- Alexa On-Demand Briefing — "Alexa, briefing" or "Alexa, mail status" triggers pyscript pipeline delivery.
- Presence Patterns — Markov transition probabilities from FP2 zone history. Predicts next zone, powers pre-activation.
- Away Patterns — Departure/return prediction with multi-trip tracking and per-person daily trip counters.
- Zone Presence / Vacancy / Preactivation — Per-zone automations: actions on occupancy, actions on vacancy, pre-stage devices when a user approaches.
- Privacy Gate — Per-person, per-tier feature suppression. T1 (intimate: bedtime, wake-up), T2 (personal: notifications, briefings), T3 (ambient: tracking, recommendations). 33 gated features, 33 per-feature overrides, hysteresis to prevent flapping.
- Coming Home — Arrival automation with pre-arrival actions and welcome scene.
- Bedtime Wind-Down — 4 detection scenarios (sleepy_tv, bed_tv, bed_idle, bed_non_sleepy) with cooldown curves, budget gate, memory integration, and privacy gating.
- Sleep Detection — FP2-based sleep state detection with configurable timeout for false positive prevention.
- Wakeup Guard — 3 variants: basic (snooze/stop), escalating (progressive volume, escape hatch), external alarm (phone alarm trigger). Mobile notification companion blueprint.
- Bedtime Routine Core — Complete shutdown sequence: TV off (CEC/script/IR fallback chain), lights, MA or Kodi media playback, countdown, bathroom guard, settling + goodnight TTS.
- Sleep Lights — Ambient lighting during sleep with configurable light and sensor pickers.
- Predictive Schedule — Calendar-aware bedtime and wake time prediction with confidence tracking.
- Music Follow-Me — Presence-based multi-room audio routing. Priority ordering, cooldown timers, multi-person detection, mixed ecosystem support (Sonos + Voice PE + Alexa). LLM-generated room-change announcements.
- Music Composer — Hybrid synthesis engine: ElevenLabs API + FluidSynth local MIDI. 11 services, staging mode for review before production, feedback mic for iterative refinement, per-agent musical identity, 9 content types (theme, chime, handoff, expertise, thinking, stinger, wake melody, bedtime, ambient).
- Music Taste — Spotify + Music Assistant playback analysis. Genre, artist, and mood trending with taste model rebuilding.
- Alexa Presence Radio — "Alexa, turn on [station]" plays on the speaker in your current room. Follow-Me interlock, multi-zone mapping, pause-vs-stop behavior, duck guard integration.
- Media Tracking — Radarr/Sonarr integration for upcoming releases and recent download notifications.
- Watch History — Kodi playback tracking with source detection (YouTube, Netflix, Prime Video, Movistar+, PVR, library) via JSON-RPC. EPG season/episode enrichment for PVR channels. Configurable minimum watch duration thresholds filter channel surfing. Channel flip counting. L2 memory logging with daily summaries. Hot context injection ("Now watching" / "Recently watched").
- Routine Fingerprinting — Greedy Markov chains from zone transition sequences. Stage tracking, ETA calculation, deviation detection with automated actions.
- Scene Learner — Learns lighting and scene preferences from user behavior. Per-zone, per-context scene storage. (Built, not yet tested.)
- User Interview — 9-category preference elicitation (identity, household, work, schedule, health, environment, media, communication, privacy). Pre-seeds from existing memory. Preferences consumed by 14 blueprints/modules for LLM prompt shaping (humor, off-limits, verbosity), sleep budget calculation (hours until wake with weekday/weekend/alt-day routing), and schedule-aware alarm/routine trigger overrides. Day-name normalization on ingest (Spanish → English). (Deployed but untested as of 2026-03-30 — live voice testing blocked until ElevenLabs subscription renews.)
- Contact History — Per-contact message logging with LLM-powered batch compression.
- Interaction Summarizer — Nightly batch compresses whisper logs via cheap LLM into digests. 3 retention modes. ~$0.001/run.
- TTS Queue — Priority queue (5 levels: emergency through ambient). Dynamic speaker discovery from zone config JSON. Ranked speaker preference, volume ducking, playback tuning, TTS completion events, caching, voice mood injection (stability + audio tag prefixes for non-agent messages).
- Duck Manager — Reference-counted volume ducking coordinator. 3 behavior modes (volume, pause, both). Snapshot/restore with user-adjustment protection. Auto-discovers Music Assistant volume aliases via entity registry — no manual config for multi-entity speakers.
- Focus Guard — Anti-ADHD nudge system. 6 nudge types (hydration, movement, screen break, meal, medication, custom). Escalating delivery, focus mode suppression, meal detection with voice interaction.
- Phone Call Detection — Defers all TTS during active phone calls.
- Refcount Bypass Watchdog — Safety net for 7+ automations sharing speaker control. Counter-based atomic operations, 2-minute polling, stale state auto-reset on HA restart.
- System Health Sensor — Validates 7 subsystems every 30 minutes: status sensors, pyscript services, helpers, pipeline entities, TTS queue, JSON configs, memory DB. Weighted scoring, automatic state transition events.
- Self-Healing Recovery — 7 recovery playbooks with exponential backoff, circuit breaker (3 retries/hour max), reload coalescing, pending-state persistence across restarts.
- Dispatcher Bypass Mode — After 3 consecutive cache failures, automatically routes to a fallback pipeline. Self-clears on recovery.
- Memory Read-Only Mode — After 3 consecutive write failures, transitions to read-only with periodic recovery probe. Reads continue normally.
- Toggle Audit Trail — Source-attributed logging of every AI kill switch change. SQLite storage, 90-day retention, critical switch alerting.
- Cross-User Memory Isolation — Owner column on L2 memory with Tier B access control. Non-owner queries see metadata but values are
[restricted]. Identity confidence gate for dual-occupancy writes.
- Calendar CRUD — Full voice-driven Google Calendar create/find/delete/edit. Recurring event support (this_instance, this_and_future, entire_series). HACS calendar_utils for UID access. Unified
calendar_event(operation=...)agent tool. - Calendar Alarm — Calendar-aware wake-up using predicted_wake_time from bedtime advisor. Workday gating, presence checking.
- Circadian Lighting — Continuous color temperature and brightness adjustment from sun elevation. Sleep mode, manual override detection, scene_learner integration.
A 6-tab, 2,521-line Lovelace dashboard for monitoring and configuring the entire AI stack:
- Overview — 13-module health glance, LLM budget gauges, per-agent cost breakdown, ElevenLabs credits, OpenRouter usage
- Configuration — 34 cards covering all subsystems, 27 kill switches, zone-to-speaker mapping, expertise routing
- User Profiles — Identity confidence scores, sleep schedule, language preferences, privacy gate with 33 per-feature overrides
- Presence — Routine tracking, predictive patterns, bedtime advisor, 8-zone FP2 presence map
- Debug — Focus guard, duck manager status, media tracking, music composer, test harness
- Memory — Recent topics, memory DB health, archive browser, relationship browser
Deep Space One — An LCARS-themed operations console panel. Very alpha. Very DS9. (Removed 2026-03-29)
Therapy mode— Deployed, untested. Live voice testing blocked until ElevenLabs subscription renews (early April 2026). See Feature Highlights above.- Deliberation mode — Multi-agent internal consensus. Multiple agents respond internally, a synthesis agent presents the unified answer or the disagreement. (Pattern 2 — not yet built.)
- Header images — 66 of 113 blueprints are missing Gemini-generated header images for the GitHub description field.
Of 113 total blueprints (79 automation + 34 script), 89 have live instances running daily. 24 have never been deployed and need community testing.
Blueprints with no deployed instance (19 automation + 6 script):
| Category | Blueprints | Why |
|---|---|---|
| Lighting/Scene | circadian_lighting, ambient_music_autoplay, scene_preference_apply, zone_preactivation |
No color_temp bulbs to test |
| Bedtime | bedtime_winddown, bedtime_last_call, bedtime_advisory_actions, calendar_alarm, wake_up_guard_external_alarm |
Not yet configured |
| Presence | away_state_actions, coming_home, routine_deviation_actions, routine_stage_actions |
Not yet configured |
| Music | music_compose_batch_trigger, music_weekly_refresh, music_assistant_follow_me_idle_off, music_compose_approve |
Not yet configured |
| Budget | budget_cost_alert |
Not yet configured |
| Voice | voice_pe_resume_media, voice_pin_action, llm_voice_script |
Not yet configured |
| Other | automation_trigger_mon, bedtime_media_play_wrapper, rickyellsplusalexa, wakeup_chime |
Utility/unused |
If you deploy any of these and they work (or don't), feedback is welcome via GitHub Issues.
- Add this repository as a HACS custom repository (category: Integration)
- Install Project Fronkensteen from HACS — pick a release, not the
mainbranch - Restart Home Assistant
- Go to Settings > Integrations > Add Integration > Project Fronkensteen
- Follow the 5-step setup wizard (feature selection, household config, speaker setup)
- Restart Home Assistant again
The installer copies all pyscript modules, packages, blueprints, helpers, and the patched ElevenLabs TTS to the correct locations. It only installs files for the feature groups you select.
Releases only. HACS installs this from a zip asset attached to each release (
project_fronkensteen.zip), so the default branch is deliberately hidden — installing frommainwould look for an asset that does not exist there and fail. Updates are then applied automatically; see Updating.
See INSTALL.md for the full 13-step manual installation guide. Read PREREQUISITES.md first for required HACS components, API keys, and hardware.
- PREREQUISITES.md — Required components, API accounts, hardware
- INSTALL.md — Step-by-step manual installation (13 steps)
- ARCHITECTURE.md — How the 38 modules, 78 blueprints, and 44 packages connect
- helpers/helpers_setup_guide.md — Helper configuration quick-start
- helpers/helpers_reference.md — Full reference for all 423 helpers
This section describes the author's hardware. See PREREQUISITES.md for minimum requirements.
Hardware:
- Raspberry Pi 5 (8GB RAM, 2TB NVMe SSD) running Home Assistant OS
- 2x Home Assistant Voice Preview Edition satellites (workshop + living room)
- 2x Aqara FP2 presence sensors (multi-zone: workshop/entrance/kitchen, living room/bedroom, bathroom shower/sink/toilet)
- Aqara G3 camera hub
- Sonos Era 100 (workshop) + Sonos Roam 2 (bathroom)
- GL-MT6000 router (OpenWrt — WiFi device tracking)
- APC UPS (backing the entire HA stack)
- Assorted Alexa devices (Rule of Acquisition #3: never spend more for an acquisition than you have to — they were already here)
Key Integrations:
- Extended OpenAI Conversation — LLM-driven conversation agents with function calling and custom personas
- Llama 4 Maverick via OpenRouter — LLM backend for all conversation agents
- ElevenLabs Custom TTS (HACS, forked at v0.6.3) — unified TTS entity with per-profile voice selection (
voice_idpassthrough), ElevenLabs v3 model support, and voice mood modulation (stability injection + audio tag prefixes). HACS auto-updates disabled. - Music Assistant — multi-room audio management
- ESPHome — device firmware and voice pipelines
- microWakeWord — on-device custom wake word detection
- Pyscript — Python scripting runtime (35 orchestration modules, 155 services)
- Serper — web search API available to all agents
- calendar_utils (HACS) — Google Calendar UID access for event CRUD operations
├── automation/ 79 automation blueprints
├── script/ 34 script blueprints
├── packages/ 44 YAML packages (sensors, automations, helpers)
├── pyscript/ 38 Python modules (168 services)
├── helpers/ 7 helper files (217 entities)
├── Extended OpenAi Conversation Prompts/
│ ├── Standard/ 5 persona prompts + shared functions
│ ├── Bedtime/ Bedtime-variant prompts
│ ├── Music Compose/ Music generation prompts
│ ├── Music Transfer/ Music playback prompts
│ └── Therapy/ Portuondo-only (deployed)
├── images/header/ 100+ Gemini-generated blueprint headers
├── readme/
│ ├── automation/ 79 automation READMEs
│ ├── script/ 33 script READMEs
│ ├── packages/ 29 package READMEs
│ └── pyscript/ 24 module READMEs
├── style-guide/ Rules of Acquisition (11 files, ~126K tokens)
└── archive/ Superseded blueprints and READMEs
The style-guide/ directory contains the Rules of Acquisition — a comprehensive YAML generation style guide used by AI agents (primarily Claude) when generating Home Assistant configurations. Named after the Ferengi code of conduct from Star Trek: Deep Space Nine, the guide enforces opinionated standards across 11 files covering core philosophy, blueprint patterns, automation patterns, conversation agents, ESPHome devices, Music Assistant integration, anti-patterns, troubleshooting, voice architecture, and QA audit checklists.
Three operational modes: BUILD (full compliance, mandatory build logs), TROUBLESHOOT (debugging focus, minimal ceremony), AUDIT (systematic violation scanning with severity classifications).
This entire configuration is vibe coded — built through conversational AI sessions with Claude (Anthropic) running inside Claude Code with custom instructions, persistent memory, and direct filesystem access to the HA config via SMB mount. Claude reads and writes files directly, commits through git, and validates configurations against the live Home Assistant instance.
No YAML in this repository was written by hand. Budget sustainability is a design principle, not an afterthought — every model choice, every fallback tier, every nightly batch job is built around the constraint of running on personal infrastructure at personal-budget costs.
| Project | Author | License | Role |
|---|---|---|---|
| Extended OpenAI Conversation | jekalmin | Apache 2.0 | Conversation agent framework — patched with 4-layer speech sanitizer that strips tool-call leaks from TTS output (Gemini bare calls, Maverick orphaned args, inline params). HACS auto-updates disabled. |
| Pyscript | Craig Barratt (@craigbarratt) | Apache 2.0 | Python scripting runtime for all 29 orchestration modules |
| Voice Assistant Long-term Memory | luuquangvu | — | Foundation of L2 memory system (SQLite+FTS5); extended with auto-relationships, scopes, tag linking, embeddings, archiving |
| ElevenLabs Custom TTS | Loryan Strant (@loryanstrant) | MIT | Voice profile system — patched with v3 mood modulation: stability slider + audio tag prefix injection for non-agent TTS. HACS auto-updates disabled. |
| microWakeWord | Kevin Ahrendt (@kahrendt) | — | On-device wake word detection for ESP32-S3; custom wake word models trained with this |
| calendar_utils | — | — | Google Calendar UID access and event deletion for voice CRUD operations |
| Project | Author | What We Used |
|---|---|---|
| Home Generative Agent | Lindo St. Angel (@goruck) | Context summarization, semantic memory search, multi-model cost tiers, tool error handling, critical action PIN flow (architectural reference — no code copied) |
| RC Home Assistant Low-VRAM | RoyalCities | Visible memory via todo list pattern, follow-up conversation state machine |
| HA Music Voice Control SpotifyPlus | brix29 | Voice control patterns for Music Assistant integration |
| Music follow-me concept | Phil Hawthorne | Presence-based audio routing inspiration |
| EL-HARP activity prediction | Academic paper | Feature taxonomy for activity prediction (time bucketing, dwell times, transition probabilities) applied to Markov presence prediction |
| arc42 architecture template | Gernot Starke, Peter Hruschka | Document structure inspiration for voice context architecture doc |
| PEveleigh dispatcher pattern | peveleigh | Basic conversation.process routing between agent instances; extended into 7-level pipeline-aware dispatcher |
| Author | Blueprints |
|---|---|
| Blackshome | Sensor light, motion-activated lighting |
| luuquangvu | Memory tool blueprints (local + full LLM) |
| TheFes | Music Assistant LLM voice script (adapted as llm_voice_script.yaml) |
| Music Assistant | Official MA blueprints |
| SpotifyPlus | SpotifyPlus integration blueprints |
| Component | Role |
|---|---|
| HA Voice Assistants (Assist Pipelines) | Centralized pipeline configuration — routing layer for all 22 persona variants |
| ElevenLabs (HACS custom + official) | Primary TTS via custom component (voice profiles + mood modulation); official integration as secondary entity source |
| Google Calendar | L3 data source for calendar promotion, alarm scheduling, and voice CRUD |
| IMAP | L3 data source for email priority filtering |
| Aqara (via ha_aqara_devices) | FP2 presence detection across 8 zones |
| Tool | Role |
|---|---|
| OpenRouter | LLM API gateway — dynamic per-model pricing, cost tracking |
| Serper | Web search API for agents |
| sqlite-vec | Semantic vector embeddings for L2 memory search |
| FluidSynth / midiutil | Local MIDI synthesis for budget-aware music composition |
| ha_text_ai | Tool-free LLM text generation for non-conversational tasks |
- Young Frankenstein (1974, Mel Brooks) — Project name ("It's pronounced Fronkensteen")
- Star Trek: Deep Space Nine (Paramount) — Quark persona, Rules of Acquisition naming, Ferengi philosophy throughout
- Rick and Morty (Adult Swim) — Rick Sanchez persona
- Deadpool (Marvel / 20th Century Studios) — Deadpool persona
- Seinfeld (NBC) — Cosmo Kramer persona
- Doctor Portuondo (Filmin series / Carlo Padial novel) — Doctor Portuondo persona
- Claude (Anthropic) — All code generation, architecture design, audit, and documentation
- Google Gemini — All header image generation (Rick & Morty / Star Trek animation styles)
To Jessica — for her patience with this obsession, for coexisting with four AI personalities she never auditioned for, and for being living proof that the best things in a smart home have nothing to do with the technology.
Trained models available at mmadalone/microwakeword:
- Hey Rick / Yo Rick
- Hey Quark / Yo Quark
This is a personal home automation configuration shared for reference and inspiration. Use whatever's useful to you. If you build on any of the blueprints, a mention would be appreciated but isn't required.
Rule of Acquisition #286: when Morn leaves, it's all over.
