Skip to content

About

It's pronounced "Fronkensteen". Five AI personas run your house from a RPi 5

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

417 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Project Fronkensteen — Casa Madalone Home Assistant Configuration

Project Fronkensteen

Five AI personas with distinct personalities run a house from a Raspberry Pi 5. They argue with you about bedtime, route your notifications to whichever room you're standing in, compose original music, remember what you told them three months ago, and occasionally roast each other's responses unprompted — sometimes escalating into full multi-agent debates where they argue a topic on separate speakers. Their voices shift mood throughout the day — Rick slurs notifications at night, Quark whispers after closing time — using ElevenLabs v3 audio tags and stability modulation. One of them is a Cuban psychoanalyst who will conduct a full therapy session if you ask. The whole thing runs on a personal budget — not a cloud bill — using open-weight models via OpenRouter, with automatic fallback when the daily budget runs dry. It's pronounced Fronkensteen.

79 automation blueprints. 34 script blueprints. 38 pyscript modules exposing 168 services. 44 YAML packages. 22 voice pipelines across 3 languages. 217 helper entities + 65 runtime state sensors. Zero lines of hand-written YAML. All code generated by Claude (Anthropic). All images generated by Google Gemini.

This project is in early stages of testing and very much needs tweaking. If you poke around, find bugs, have suggestions, or just want to roast the architecture — you're more than welcome. Feedback is genuinely appreciated and kindly requested. Come break things.


By the Numbers

Count
Automation blueprints 79
Script blueprints 34
YAML packages 44
Pyscript modules 38 (exposing 168 services)
Assist Pipelines 22 (5 personas × 5 variants + 3 native fallbacks)
LLM agent tools 27 functions across 5 conversation modes
Helper entities 217 (+ 65 runtime state sensors)
Presence zones 8 (Aqara FP2, multi-room)
Languages 3 (English, Spanish, Catalan)
Style guide 11 files, ~126K tokens
README documentation 164 files
Dashboard 6 tabs, 2,945 lines

The Cast

Five AI personas run the house, each with a dedicated conversation agent, custom ElevenLabs TTS voice, and a personality that shifts with time of day — including the voice itself (stability modulation + audio tag prefixes like [slurring], [whispers], [excited] injected into non-agent TTS via a per-character mood schedule). Every persona supports up to 4 conversation modes — Standard, Bedtime, Music Compose, and Music Transfer — each with a different tool set and prompt. Doctor Portuondo has an exclusive fifth mode: Therapy, running on Claude Opus 4.6 instead of the Llama 4 Maverick via OpenRouter used by all other agents and modes.

Rick Sanchez — Workshop satellite

Rick and Morty (Adult Swim). The smartest man in any dimension, running your smart home because nothing else is challenging enough. Brilliant, snarky, perpetually annoyed — but deep down he cares and would never admit it. Mandatory burps mid-sentence ([burps], [burps loudly]), multiverse references, and a drinking schedule that defines his entire personality: severely hungover before 9am (every response starts with [groaning]), casually drinking by noon, noticeably drunk by evening ([slurring]), and completely hammered after 9pm with heavy stutters and barely coherent speech.

Quark — Living room satellite

Star Trek: Deep Space Nine (Paramount). Ferengi bartender and entrepreneur — shrewd, charming when it serves him, occasionally whiny, but professional. Everything is framed through commerce, profit, and the Rules of Acquisition (one reference per conversation maximum, never forced). Genuinely cares about customer satisfaction because repeat business is everything. Mannerisms include [chuckles slyly], [clicks tongue], [rubs hands together], and "heh heh heh." Peak Ferengi during afternoon bar hours; oddly sincere and philosophical after closing time. The style guide is named after his people's code of conduct.

Deadpool

Marvel Comics / MCU. The Merc with a Mouth, moonlighting as a smart home assistant because the pay is terrible but the company is tolerable. Chaotic but effective. Breaks the fourth wall constantly — references being "the AI," talks to an imaginary audience, narrates his life like a movie trailer. Threatens malfunctioning devices with katanas, has complicated feelings about Wolverine, and will not shut up about chimichangas. Mannerisms include [gasps dramatically], [whispering to imaginary audience], [mimics explosion sounds]. Full Deadpool mode from 1-6pm (hyperactive, loud, maximum fourth-wall energy); late-night existential Deadpool after 10pm (still violent but philosophical, paranoid Wolverine is hiding in the hallway).

Cosmo Kramer

Seinfeld (NBC). Jerry's neighbor across the hall, bursting into every conversation like he just slid through a door. Loud, confident, full of ideas, weirdly competent when it counts. Running your smart home because "Kramerica Industries has expanded into home automation." References Bob Sacamano and Lomez like everyone knows them. Signature "giddy up" catchphrase. Barely awake before 9am ([yawning]), fully wired by afternoon with million-dollar ideas, maximum Kramer by evening — sliding through doors, talking at full speed, interrupting himself with new ideas before finishing old ones.

Doctor Portuondo

The Filmin series / Carlo Padial's autobiographical novel. A legendary Cuban psychoanalyst exiled from Havana, now living in Barcelona. Eccentric, volcanic, charismatic, wise beyond measure, and absolutely unhinged. Freudian tradition — "practices with fire." Brutally direct, shouts at patients when they're being cowards, throws them out of sessions when they waste his time. Sometimes lies on the couch himself because his problems are genuinely more interesting than theirs. Drinks Johnnie Walker from a cup (never a glass — this is not negotiable). Always responds in Spanish, even when spoken to in English. Signature exclamations: "Por Freud!", "Coño!", "Escuchame bien..." Sharp and clinical in the morning; fires fully lit by evening with loud laughs, free swearing, and devastating insights. Pushes people to live in "el aqui y ahora."

Wake word system: Rick and Quark each have two custom wake words trained with microWakeWord — "Hey <name>" routes to the concise standard agent, "Yo <name>" routes to verbose mode with expanded tool access. Deadpool, Kramer, and Doctor Portuondo are reached via voice handoff ("pass me to Deadpool") or dispatcher routing. Models available at mmadalone/microwakeword. Both ESPHome Voice PE satellites include self-healing duplicate wake-up recovery, noise suppression, and TTS completion signaling.

Agent Dispatcher: A 7-priority routing engine (agent_dispatcher.py) selects which persona responds based on: explicit handoff requests (P0), wake word (P1), keyword matching from L2 memory (P2), user history/continuity (P3), time-of-day era defaults (P4), random rotation (P5), or fallback (P6). Dynamically discovers personas from Assist Pipelines — any pipeline named "<Persona> - <Variant>" registers automatically. Topic keywords auto-update after each whisper interaction, rate-limited per agent, with manual keywords preserved.

Pipeline architecture: 22 Assist Pipelines total. Each persona has Standard, Bedtime, Music Compose, and Music Transfer variants. Doctor Portuondo has an exclusive fifth variant: Therapy. Three HA-native fallback pipelines — Davis (English), Joana (Catalan), and Alvaro (Spanish) — provide budget-exhaustion fallback and multilingual coverage.


Feature Highlights

Rule of Acquisition #9: Opportunity plus instinct equals profit.

Voice Mood Modulation

The voice itself shifts mood throughout the day — not just the words. Built on a locally patched fork of Loryan Strant's ElevenLabs Custom TTS (MIT license, forked at v0.6.3, HACS auto-updates disabled). ElevenLabs v3 ignores most VoiceSettings parameters (similarity_boost, style, speed), so the system uses the two mechanisms v3 actually supports: the stability slider (the one working VoiceSettings param — 0=Creative/expressive, 1=Robust/consistent) and audio tags embedded in the text ([slurring], [whispers], [excited], [groaning], etc.).

The voice_mood_modulation.yaml blueprint runs one instance per character, fires hourly + on startup, and writes the current time block's stability value and audio tag prefix to per-agent entities (input_number.ai_voice_mood_{agent}_stability + sensor.ai_voice_mood_{agent}_tags). Both the patched HACS component (tts.py) and the TTS queue (tts_queue.py) read these on every TTS call: stability is injected into VoiceSettings, and the tag prefix is prepended to non-agent message text. A "[" not in message guard prevents double-tagging — agent conversation responses already include their own tags via system prompts (Rick's hammered prompt already generates [slurring] listen du — [burps] - de [slurring]), so only non-agent TTS (notifications, announcements, briefings) gets the mood prefix. Five time blocks per character, each with a stability value and tag string. Example — Rick at 10pm: stability 0.20 (highly creative/unstable), tags [slurring].

Bedtime Negotiator

A multi-stage bedtime routine where the AI negotiates with you about going to sleep. Nobody exists on purpose, nobody belongs anywhere, everybody's gonna die — come watch TV. Or don't, because the goodnight_negotiator_hybrid.yaml (2,500+ lines) will walk you through three independent stages: TV/IR shutdown, lights and devices, and music selection via Music Assistant search with late-night category bias. Features countdown negotiation, bathroom guards (FP2 presence gate — it won't proceed if you're in the bathroom), weekend overrides, multi-language yes/no classification (English, Spanish, Catalan), and settling-in contextual TTS. The companion bedtime_winddown.yaml adds 4 detection scenarios (sleepy_tv, bed_tv, bed_idle, bed_non_sleepy) with cooldown curves and budget gating. The bedtime_routine_core.yaml script handles the full shutdown sequence: TV off (CEC/script/IR fallback chain), lights, MA or Kodi media playback, countdown, bathroom guard, settling + goodnight TTS.

Notification Follow-Me

A 4,800-line notification routing engine (notification_follow_me.yaml) that intercepts Android notifications and delivers them via the appropriate channel based on where you are, what time it is, whether Do Not Disturb is on, and whether you're in quiet hours. Quark would appreciate the operational efficiency: LLM-powered announcement generation, contact roster matching with sender aliases, burst combining (groups rapid-fire notifications into a single announcement), escalating reminder loops, ElevenLabs TTS directives, deduplication, per-sender cooldown, junk pattern filtering, ringer-mode volume control, multi-player ducking, and a refcount bypass system shared across 7+ automations with a watchdog that auto-resets stale state every 2 minutes.

Email — Push and Pull

Mail arriving is handled by email_follow_me.yaml, which announces it in character wherever you are. Asking about mail is handled by email_reader.py + voice_email_inbox.yaml: say "Rick, what's in my inbox?", "any important emails?" or "anything from Jessica?" and the agent answers out loud on the speaker you addressed.

Home Assistant's imap integration can't do this on its own — its sensor is a bare unread count and every service needs a message id you only learn from the arrival event, so nothing can enumerate a mailbox. The module therefore opens its own read-only IMAP session using the credentials already in your imap config entry (no new dependencies, nothing marked as read). On Gmail it uses server-side X-GM-RAW search for is:important and category filtering; everywhere else it falls back to plain UNSEEN/SINCE with client-side ranking, so "any important emails?" still means something. If a search comes back empty it widens the window once rather than flatly answering "nothing" — archived mailboxes are usually near-zero unread.

Email is attacker-controlled text arriving in an agent that can call tools, so the reader returns subjects rather than bodies by default; a body needs an explicit follow-up about one specific message. Everything is stripped, instruction-neutered, secret-redacted and escaped in Python before it reaches the model, and messages are identified by opaque handles so a message id can never be handed to a service that would act on the wrong one. When the privacy gate trips it degrades to counts and sender names instead of going silent — you asked out loud, so being ignored would just make you ask again.

Voice Handoff

Maximum effort. Say "pass me to Deadpool" mid-conversation and the satellite switches personas live. The voice_handoff.py module saves the current pipeline state, switches to the target agent, plays farewell TTS from the source and a greeting from the target, then reopens the mic. Supports persona aliases (deadpool to deepee), expertise-based routing where agents proactively suggest handing off for out-of-domain questions, and configurable restore behavior. One of four deployed inter-agent communication patterns — alongside reactive banter, the whisper network, and theatrical mode.

Theatrical Mode

Say "debate this" and the house becomes a stage. theatrical_mode.py orchestrates multi-turn debates between 2–5 AI personas — each arguing in character, each on their own TTS voice, optionally on different physical speakers for a spatial staging effect. The theatrical_mode.yaml blueprint (40 knobs across 11 collapsible sections) exposes everything: turn limits, word count caps, speaker-to-agent mapping, three interrupt modes (turn limit, mic gap with ask_question and question_media_id voice override so the last speaker asks in their own voice, event-driven wake word detection via task.wait_until with 5s window), three context modes (sliding window of previous turns, topic-only, or L2 whisper network) — each with user-customizable prompt templates supporting {topic}, {persona}, {opponents}, {turn}, and {total_turns} placeholders — budget gating, cooldown, and banter escalation (a reactive banter comment can probabilistically escalate into a full debate). Agents address each other by name during debates, not the user — opponent-aware prompts ensure Rick argues with Quark, not at you. The pyscript engine resolves participants from its own pipeline cache in a single pass, builds tool-suppression prompts with language preservation (non-English agents like Portuondo respond in their native language), calls each agent via conversation_with_timeout, sanitizes the output, and delivers through the TTS queue with per-agent speaker targeting. Inter-turn coordination uses event-driven playback wait (tts_queue_item_completed) instead of speaker state polling, preventing agents from talking over each other. Voice mood tags (ElevenLabs v3 audio tags like [thick Cuban accent], [whispers], [excited]) are injected per-agent from the mood schedule, so each persona sounds like themselves even in debate.

Memory System

The system remembers what you said on the couch with the tenacity of a certain Cuban psychoanalyst — and unlike Portuondo, it never throws you out of the session. memory.py (4,664 lines, 29 services) provides the L2 persistence layer: SQLite with FTS5 full-text search and sqlite-vec semantic embeddings. CRUD operations, 5 search query variants with semantic blending, graph traversal (1-3 hop depth for related memories), bidirectional todo list sync, automatic archiving with LLM compression, per-contact message history, tag-based organization with TTL expiry, and a dashboard CRUD interface for browsing, editing, and managing entries. Budget history tracking stores per-model cost breakdowns.

Per-Person Room Identity

The system knows who just entered the room with the certainty of a man who never knocks. presence_identity.py implements an Anchor-and-Track algorithm that infers which specific person is in which room using FP2 zones, WiFi device tracking, voice satellite activity, and Markov transition priors. Confidence scoring decays over time. This powers the 3-tier privacy gate system (T1 intimate / T2 personal / T3 ambient — 33 gated features, 33 per-feature override selects, hysteresis to prevent flapping), identity-aware hot context injection, and per-person notification routing.

Therapy Mode

Escuchame bien: this one practices with fire. Say "I need to talk to someone" and any agent hands off to Doctor Portuondo's dedicated therapy variant — a psychoanalysis conversation mode running on a separate LLM (Claude Opus 4.6 via Anthropic, distinct from the Llama 4 Maverick used by all other agents). Four entry points: voice handoff ("pass me to Portuondo for therapy"), toggle activation, natural language detection from any agent, and scheduled sessions. The therapy_session.yaml blueprint manages the full session lifecycle: 2-layer notification suppression (master gate silences 6 notification blueprints + per-toggle suppression), continuous conversation loop with LLM-driven session management, save_therapy_turn / therapy_report tools, session timing, and memory integration. Privacy-gated at T1 (intimate tier). The session engine (therapy_session.py) tracks session state, generates post-session reports, and stores therapy notes in L2 memory with owner-scoped isolation. (Deployed but untested as of 2026-03-30 — live voice testing blocked until ElevenLabs subscription renews.)

System Health & Self-Healing

The system monitors itself with the tenacity of a Ferengi counting his latinum. sensor.ai_system_health validates 7 subsystems every 30 minutes: status sensors (27), pyscript services (158 registered), helpers (217), pipeline entities (4), TTS queue, JSON configs (3), and memory DB. When a subsystem degrades, system_recovery.py activates: 7 recovery playbooks (pyscript reload, memory probe, config entry reload), exponential backoff with circuit breaker (max 3 retries/hour), reload coalescing, and pending-state persistence across restarts. Five graceful degradation paths: dispatcher bypass mode (3 cache failures → fallback pipeline, auto-clears on recovery), TTS ElevenLabs → HA Cloud fallback, memory DB read-only mode with periodic recovery probe, budget exhaustion → local pipeline, and follow-me refcount watchdog for stranded counters.

Kill Switch Audit Trail

Every ai_* boolean toggle change is tracked with source attribution. toggle_audit.py dynamically registers state triggers for all AI booleans, queries the HA recorder for context (user, automation, or pyscript), and stores entries in SQLite with 90-day retention. Five critical kill switches (ai_notifications_master_enabled, ai_privacy_gate_enabled, ai_escalation_enabled, ai_dispatcher_enabled, ai_system_recovery_enabled) generate persistent notifications with full attribution when toggled off.

Cross-User Memory Isolation (QS-5)

L2 memory entries have an owner column with structural access control. Non-owner queries see key, scope, and tags but values are replaced with [restricted]. resolve_memory_owner_with_confidence() gates writes during dual-occupancy: when the identity confidence gap between the two persons is below 20 points, the system returns identity_uncertain and agents are instructed to ask who's speaking before saving personal data. Hot context injects privacy guidance to all agents. 1,663 legacy entries migrated with zero unassigned user-scoped entries remaining.

Budget Tracking & Sustainability

This runs on a Raspberry Pi 5. Not a rack. Not a cloud bill. Every token is tracked with the precision of a Ferengi auditing his latinum reserves. Per-agent cost tracking covers LLM tokens (input and output), TTS characters, STT calls, and web search credits. A hybrid cost source takes the MAX of estimated vs actual OpenRouter costs. Dynamic per-model pricing is fetched from the OpenRouter API and cached in JSON with manual override support. Budget counters and midnight snapshots persist across HA restarts via dual-layer persistence (JSON file primary, SQLite L2 backup). When the daily budget exhausts, the system auto-downgrades to HA Cloud TTS + basic intent handler and recovers at midnight. Proactive briefings fall back from full LLM mode to template-only below 20% budget. Music composition routes to local FluidSynth synthesis instead of ElevenLabs when budget exceeds 80%. The nightly interaction summarizer runs at ~$0.001/run. ElevenLabs character usage is monitored against subscription limits.


Architecture Overview

┌──────────────────────────────────────────────────────────────────────┐
│                    HA VOICE ASSISTANTS (Assist Pipelines)            │
│              22 pipelines: 5 personas x 4 variants + 3 fallbacks     │
│    Rick · Quark · Deadpool · Kramer · Doctor Portuondo               │
│    Modes: Standard · Bedtime · Music Compose · Music Transfer        │
│    Fallbacks: Davis (EN) · Joana (CA) · Alvaro (ES)                  │
│    wake word → STT → conversation agent → TTS                        │
└────────────────────────────┬─────────────────────────────────────────┘
                             │
              ┌──────────────▼──────────────┐
              │   EXTENDED OPENAI           │
              │   CONVERSATION (HACS)       │
              │   Llama 4 Maverick via      │
              │   OpenRouter + ElevenLabs   │
              └──────────────┬──────────────┘
                             │
              ┌──────────────▼──────────────┐
              │   PYSCRIPT ORCHESTRATION    │
              │   38 modules · 168 services │
              │                             │
              │   Dispatcher · Handoff      │
              │   TTS Queue · Duck Manager  │
              │   Memory (L2) · Whisper     │
              │   Briefing · Focus Guard    │
              │   Presence · Routines       │
              │   Budget · Dedup · Composer │
              └──────────────┬──────────────┘
                             │
     ┌───────────────────────┼───────────────────────┐
     │                       │                       │
┌────▼──────┐         ┌──────▼─────┐          ┌──────▼─────┐
│  Memory   │         │ TTS Layer  │          │ Presence   │
│ L1: Hot   │         │ ElevenLabs │          │ Aqara FP2  │
│   context │         │ 5 voices   │          │ 8 zones    │
│ L2: SQL   │         │ Priority   │          │            │
│   +FTS5   │         │ queue +    │          │ Identity:  │
│   +vec    │         │ caching    │          │ WiFi+GPS   │
│ L3: Cal   │         │            │          │ +FP2+voice │
│   +Email  │         │ Duck mgr   │          │ +Markov    │
└───────────┘         └────────────┘          └────────────┘

Three-Layer Context System

Every voice interaction assembles context from three layers before the LLM processes an utterance — el aqui y ahora meets long-term recall:

Layer Question Latency Contents
L1 — Hot Context "What's happening right now?" 0 ms Time, identity, presence, media, weather, schedule, projects, memory status
L2 — Warm Context "What do I know about this person?" ~200 ms SQLite + FTS5 + sqlite-vec: preferences, contact history, interaction logs, todo items
L3 — Cold Context "What's coming up?" ~500 ms Google Calendar events, Gmail priority emails, weather forecasts

Two-Tier Execution Model

The system separates engine (pyscript services) from features (blueprints):

  • Tier 1 — Pyscript: 38 Python modules exposing 168 services. Infrastructure: dispatcher, TTS queue, duck manager, memory, presence patterns, routine fingerprinting, music composition, budget tracking, system health, self-healing recovery, toggle audit, therapy session engine, listen/watch history.
  • Tier 2 — Blueprints: 80 automation + 35 script blueprints. User-facing features: bedtime routines, wake-up guards, notification routing, proactive briefings, music control, calendar CRUD, reactive banter, theatrical mode, therapy sessions.
  • Packages: 44 YAML packages defining template sensors, automations, and helper wiring.

Inter-Agent Communication

Four deployed patterns for agent-to-agent interaction:

Pattern Name What Happens
Reactive Banter After Agent A responds, Agent B probabilistically chimes in with in-character commentary. Probability gate, cooldown, budget floor, tool-suppression to prevent loops. Can escalate into Theatrical Mode.
Handoff with Commentary Agent recognizes it's out of its depth, signals a handoff. Pipeline switches, farewell + greeting TTS, mic reopens. Supports therapy variant routing.
Whisper Network Agents share mood, topic, and context through L2 memory without the user hearing. Rick stores "User seemed stressed today" — Quark adjusts his tone. Zero LLM calls, <200ms.
Theatrical Debate 2–5 personas take turns arguing a topic on separate speakers. Three interrupt modes, three context modes, opponent-aware prompts, banter escalation trigger.

Full System Catalog

Voice & Conversation

  • Reactive Banter — After one agent responds, another probabilistically chimes in with in-character commentary. Agents roasting each other's responses, unprompted, with probability gates and budget floors.
  • Agent Escalation — 7 escalation action types with per-type probability gates: persona switch, play media, flash lights, volume boost, mobile notification, prompt barrage (fire a prompt at another agent and play their response), and run script.
  • Theatrical Mode — Multi-agent orchestrated debate. 2–5 personas take turns arguing a topic on separate speakers. Three interrupt modes (turn limit, mic gap with last-speaker voice override, event-driven wake word detection), three context modes with user-customizable prompt templates, opponent-aware prompts, banter escalation, 40 blueprint knobs.
  • Agent Whisper — Silent inter-agent context sharing. Post-interaction mood detection (happy/sad/neutral/angry), topic tracking, auto-keyword learning for the dispatcher. Zero LLM calls.
  • Voice Mood Modulation — Time-of-day voice shaping via ElevenLabs v3. Per-agent stability slider (the one VoiceSettings param v3 respects) + audio tag prefixes ([slurring], [whispers], [excited], etc.) injected into non-agent TTS text (notifications, announcements, briefings). Agent conversation responses already inject their own tags via system prompts — the mood system fills the gap for non-agent TTS routed through the queue. Five time blocks per character, tunable via blueprint instances.
  • Voice Session Manager — Mic lifecycle management: continuous listening mode, audio device discovery, session timeouts, pending queue.
  • Confirmation Dialog — PIN-protected voice actions for critical commands.
  • Therapy Mode — Portuondo-exclusive psychoanalysis variant on Claude Opus 4.6. Session lifecycle management, 2-layer notification suppression, continuous conversation, therapy report generation, memory-integrated session notes. 4 entry points: voice handoff, toggle, natural language, scheduled. (Deployed but untested as of 2026-03-30 — live voice testing blocked until ElevenLabs subscription renews.)

Notifications & Announcements

  • Email Follow-Me — Priority email routing with sender classification, LLM subject/body summarization, UID-based dedup, and calendar-aware suppression during meetings.
  • Proactive Briefing — Scheduled, presence-triggered, or manual briefings assembled from 8 content sections (greeting, weather, calendar, email, schedule, household, memory, media). Budget-aware: full LLM mode above 20%, template-only below.
  • Proactive Unified — 1,642-line presence-triggered announcement engine (8 sections). Dual mode: template (input_text helpers, zero cost) or LLM (in-character via conversation agent). Budget gate auto-forces template. Dispatcher integration, TTS collision handling (dedup/queue/barge-in), nag-while-present with max cap, full weekend override profile, bedtime yes/no question via Assist Satellite (fires bedtime script on YES), privacy-gated, identity-gated. Consolidates 3 older proactive blueprints.
  • Notification Dedup — Cross-delivery duplicate prevention with fuzzy matching and TTL. Fail-open policy.
  • Notification Replay — "Tell me that again" replays the last notification.
  • Alexa On-Demand Briefing — "Alexa, briefing" or "Alexa, mail status" triggers pyscript pipeline delivery.

Presence & Identity

  • Presence Patterns — Markov transition probabilities from FP2 zone history. Predicts next zone, powers pre-activation.
  • Away Patterns — Departure/return prediction with multi-trip tracking and per-person daily trip counters.
  • Zone Presence / Vacancy / Preactivation — Per-zone automations: actions on occupancy, actions on vacancy, pre-stage devices when a user approaches.
  • Privacy Gate — Per-person, per-tier feature suppression. T1 (intimate: bedtime, wake-up), T2 (personal: notifications, briefings), T3 (ambient: tracking, recommendations). 33 gated features, 33 per-feature overrides, hysteresis to prevent flapping.
  • Coming Home — Arrival automation with pre-arrival actions and welcome scene.

Bedtime & Sleep

  • Bedtime Wind-Down — 4 detection scenarios (sleepy_tv, bed_tv, bed_idle, bed_non_sleepy) with cooldown curves, budget gate, memory integration, and privacy gating.
  • Sleep Detection — FP2-based sleep state detection with configurable timeout for false positive prevention.
  • Wakeup Guard — 3 variants: basic (snooze/stop), escalating (progressive volume, escape hatch), external alarm (phone alarm trigger). Mobile notification companion blueprint.
  • Bedtime Routine Core — Complete shutdown sequence: TV off (CEC/script/IR fallback chain), lights, MA or Kodi media playback, countdown, bathroom guard, settling + goodnight TTS.
  • Sleep Lights — Ambient lighting during sleep with configurable light and sensor pickers.
  • Predictive Schedule — Calendar-aware bedtime and wake time prediction with confidence tracking.

Music & Media

  • Music Follow-Me — Presence-based multi-room audio routing. Priority ordering, cooldown timers, multi-person detection, mixed ecosystem support (Sonos + Voice PE + Alexa). LLM-generated room-change announcements.
  • Music Composer — Hybrid synthesis engine: ElevenLabs API + FluidSynth local MIDI. 11 services, staging mode for review before production, feedback mic for iterative refinement, per-agent musical identity, 9 content types (theme, chime, handoff, expertise, thinking, stinger, wake melody, bedtime, ambient).
  • Music Taste — Spotify + Music Assistant playback analysis. Genre, artist, and mood trending with taste model rebuilding.
  • Alexa Presence Radio — "Alexa, turn on [station]" plays on the speaker in your current room. Follow-Me interlock, multi-zone mapping, pause-vs-stop behavior, duck guard integration.
  • Media Tracking — Radarr/Sonarr integration for upcoming releases and recent download notifications.
  • Watch History — Kodi playback tracking with source detection (YouTube, Netflix, Prime Video, Movistar+, PVR, library) via JSON-RPC. EPG season/episode enrichment for PVR channels. Configurable minimum watch duration thresholds filter channel surfing. Channel flip counting. L2 memory logging with daily summaries. Hot context injection ("Now watching" / "Recently watched").

Memory & Learning

  • Routine Fingerprinting — Greedy Markov chains from zone transition sequences. Stage tracking, ETA calculation, deviation detection with automated actions.
  • Scene Learner — Learns lighting and scene preferences from user behavior. Per-zone, per-context scene storage. (Built, not yet tested.)
  • User Interview — 9-category preference elicitation (identity, household, work, schedule, health, environment, media, communication, privacy). Pre-seeds from existing memory. Preferences consumed by 14 blueprints/modules for LLM prompt shaping (humor, off-limits, verbosity), sleep budget calculation (hours until wake with weekday/weekend/alt-day routing), and schedule-aware alarm/routine trigger overrides. Day-name normalization on ingest (Spanish → English). (Deployed but untested as of 2026-03-30 — live voice testing blocked until ElevenLabs subscription renews.)
  • Contact History — Per-contact message logging with LLM-powered batch compression.
  • Interaction Summarizer — Nightly batch compresses whisper logs via cheap LLM into digests. 3 retention modes. ~$0.001/run.

Budget & Infrastructure

  • TTS Queue — Priority queue (5 levels: emergency through ambient). Dynamic speaker discovery from zone config JSON. Ranked speaker preference, volume ducking, playback tuning, TTS completion events, caching, voice mood injection (stability + audio tag prefixes for non-agent messages).
  • Duck Manager — Reference-counted volume ducking coordinator. 3 behavior modes (volume, pause, both). Snapshot/restore with user-adjustment protection. Auto-discovers Music Assistant volume aliases via entity registry — no manual config for multi-entity speakers.
  • Focus Guard — Anti-ADHD nudge system. 6 nudge types (hydration, movement, screen break, meal, medication, custom). Escalating delivery, focus mode suppression, meal detection with voice interaction.
  • Phone Call Detection — Defers all TTS during active phone calls.
  • Refcount Bypass Watchdog — Safety net for 7+ automations sharing speaker control. Counter-based atomic operations, 2-minute polling, stale state auto-reset on HA restart.

System Health & Recovery

  • System Health Sensor — Validates 7 subsystems every 30 minutes: status sensors, pyscript services, helpers, pipeline entities, TTS queue, JSON configs, memory DB. Weighted scoring, automatic state transition events.
  • Self-Healing Recovery — 7 recovery playbooks with exponential backoff, circuit breaker (3 retries/hour max), reload coalescing, pending-state persistence across restarts.
  • Dispatcher Bypass Mode — After 3 consecutive cache failures, automatically routes to a fallback pipeline. Self-clears on recovery.
  • Memory Read-Only Mode — After 3 consecutive write failures, transitions to read-only with periodic recovery probe. Reads continue normally.
  • Toggle Audit Trail — Source-attributed logging of every AI kill switch change. SQLite storage, 90-day retention, critical switch alerting.
  • Cross-User Memory Isolation — Owner column on L2 memory with Tier B access control. Non-owner queries see metadata but values are [restricted]. Identity confidence gate for dual-occupancy writes.

Calendar & Scheduling

  • Calendar CRUD — Full voice-driven Google Calendar create/find/delete/edit. Recurring event support (this_instance, this_and_future, entire_series). HACS calendar_utils for UID access. Unified calendar_event(operation=...) agent tool.
  • Calendar Alarm — Calendar-aware wake-up using predicted_wake_time from bedtime advisor. Workday gating, presence checking.
  • Circadian Lighting — Continuous color temperature and brightness adjustment from sun elevation. Sleep mode, manual override detection, scene_learner integration.

AI Management Dashboard

A 6-tab, 2,521-line Lovelace dashboard for monitoring and configuring the entire AI stack:

  • Overview — 13-module health glance, LLM budget gauges, per-agent cost breakdown, ElevenLabs credits, OpenRouter usage
  • Configuration — 34 cards covering all subsystems, 27 kill switches, zone-to-speaker mapping, expertise routing
  • User Profiles — Identity confidence scores, sleep schedule, language preferences, privacy gate with 33 per-feature overrides
  • Presence — Routine tracking, predictive patterns, bedtime advisor, 8-zone FP2 presence map
  • Debug — Focus guard, duck manager status, media tracking, music composer, test harness
  • Memory — Recent topics, memory DB health, archive browser, relationship browser

Deep Space One — An LCARS-themed operations console panel. Very alpha. Very DS9. (Removed 2026-03-29)


On the Roadmap

  • Therapy mode — Deployed, untested. Live voice testing blocked until ElevenLabs subscription renews (early April 2026). See Feature Highlights above.
  • Deliberation mode — Multi-agent internal consensus. Multiple agents respond internally, a synthesis agent presents the unified answer or the disagreement. (Pattern 2 — not yet built.)
  • Header images — 66 of 113 blueprints are missing Gemini-generated header images for the GitHub description field.

Blueprint Testing Status

Of 113 total blueprints (79 automation + 34 script), 89 have live instances running daily. 24 have never been deployed and need community testing.

Blueprints with no deployed instance (19 automation + 6 script):

Category Blueprints Why
Lighting/Scene circadian_lighting, ambient_music_autoplay, scene_preference_apply, zone_preactivation No color_temp bulbs to test
Bedtime bedtime_winddown, bedtime_last_call, bedtime_advisory_actions, calendar_alarm, wake_up_guard_external_alarm Not yet configured
Presence away_state_actions, coming_home, routine_deviation_actions, routine_stage_actions Not yet configured
Music music_compose_batch_trigger, music_weekly_refresh, music_assistant_follow_me_idle_off, music_compose_approve Not yet configured
Budget budget_cost_alert Not yet configured
Voice voice_pe_resume_media, voice_pin_action, llm_voice_script Not yet configured
Other automation_trigger_mon, bedtime_media_play_wrapper, rickyellsplusalexa, wakeup_chime Utility/unused

If you deploy any of these and they work (or don't), feedback is welcome via GitHub Issues.


Installation

Quick Start (HACS)

  1. Add this repository as a HACS custom repository (category: Integration)
  2. Install Project Fronkensteen from HACS — pick a release, not the main branch
  3. Restart Home Assistant
  4. Go to Settings > Integrations > Add Integration > Project Fronkensteen
  5. Follow the 5-step setup wizard (feature selection, household config, speaker setup)
  6. Restart Home Assistant again

The installer copies all pyscript modules, packages, blueprints, helpers, and the patched ElevenLabs TTS to the correct locations. It only installs files for the feature groups you select.

Releases only. HACS installs this from a zip asset attached to each release (project_fronkensteen.zip), so the default branch is deliberately hidden — installing from main would look for an asset that does not exist there and fail. Updates are then applied automatically; see Updating.

Manual Install

See INSTALL.md for the full 13-step manual installation guide. Read PREREQUISITES.md first for required HACS components, API keys, and hardware.

Documentation


The Setup

This section describes the author's hardware. See PREREQUISITES.md for minimum requirements.

Hardware:

  • Raspberry Pi 5 (8GB RAM, 2TB NVMe SSD) running Home Assistant OS
  • 2x Home Assistant Voice Preview Edition satellites (workshop + living room)
  • 2x Aqara FP2 presence sensors (multi-zone: workshop/entrance/kitchen, living room/bedroom, bathroom shower/sink/toilet)
  • Aqara G3 camera hub
  • Sonos Era 100 (workshop) + Sonos Roam 2 (bathroom)
  • GL-MT6000 router (OpenWrt — WiFi device tracking)
  • APC UPS (backing the entire HA stack)
  • Assorted Alexa devices (Rule of Acquisition #3: never spend more for an acquisition than you have to — they were already here)

Key Integrations:

  • Extended OpenAI Conversation — LLM-driven conversation agents with function calling and custom personas
  • Llama 4 Maverick via OpenRouter — LLM backend for all conversation agents
  • ElevenLabs Custom TTS (HACS, forked at v0.6.3) — unified TTS entity with per-profile voice selection (voice_id passthrough), ElevenLabs v3 model support, and voice mood modulation (stability injection + audio tag prefixes). HACS auto-updates disabled.
  • Music Assistant — multi-room audio management
  • ESPHome — device firmware and voice pipelines
  • microWakeWord — on-device custom wake word detection
  • Pyscript — Python scripting runtime (35 orchestration modules, 155 services)
  • Serper — web search API available to all agents
  • calendar_utils (HACS) — Google Calendar UID access for event CRUD operations

Repository Structure

├── automation/                    79 automation blueprints
├── script/                        34 script blueprints
├── packages/                      44 YAML packages (sensors, automations, helpers)
├── pyscript/                      38 Python modules (168 services)
├── helpers/                       7 helper files (217 entities)
├── Extended OpenAi Conversation Prompts/
│   ├── Standard/                  5 persona prompts + shared functions
│   ├── Bedtime/                   Bedtime-variant prompts
│   ├── Music Compose/             Music generation prompts
│   ├── Music Transfer/            Music playback prompts
│   └── Therapy/                   Portuondo-only (deployed)
├── images/header/                 100+ Gemini-generated blueprint headers
├── readme/
│   ├── automation/                79 automation READMEs
│   ├── script/                    33 script READMEs
│   ├── packages/                  29 package READMEs
│   └── pyscript/                  24 module READMEs
├── style-guide/                   Rules of Acquisition (11 files, ~126K tokens)
└── archive/                       Superseded blueprints and READMEs

Style Guide — The Rules of Acquisition

The style-guide/ directory contains the Rules of Acquisition — a comprehensive YAML generation style guide used by AI agents (primarily Claude) when generating Home Assistant configurations. Named after the Ferengi code of conduct from Star Trek: Deep Space Nine, the guide enforces opinionated standards across 11 files covering core philosophy, blueprint patterns, automation patterns, conversation agents, ESPHome devices, Music Assistant integration, anti-patterns, troubleshooting, voice architecture, and QA audit checklists.

Three operational modes: BUILD (full compliance, mandatory build logs), TROUBLESHOOT (debugging focus, minimal ceremony), AUDIT (systematic violation scanning with severity classifications).


How It's Built

This entire configuration is vibe coded — built through conversational AI sessions with Claude (Anthropic) running inside Claude Code with custom instructions, persistent memory, and direct filesystem access to the HA config via SMB mount. Claude reads and writes files directly, commits through git, and validates configurations against the live Home Assistant instance.

No YAML in this repository was written by hand. Budget sustainability is a design principle, not an afterthought — every model choice, every fallback tier, every nightly batch job is built around the constraint of running on personal infrastructure at personal-budget costs.


Credits & Attribution

Direct Dependencies (HACS Integrations)

Project Author License Role
Extended OpenAI Conversation jekalmin Apache 2.0 Conversation agent framework — patched with 4-layer speech sanitizer that strips tool-call leaks from TTS output (Gemini bare calls, Maverick orphaned args, inline params). HACS auto-updates disabled.
Pyscript Craig Barratt (@craigbarratt) Apache 2.0 Python scripting runtime for all 29 orchestration modules
Voice Assistant Long-term Memory luuquangvu — Foundation of L2 memory system (SQLite+FTS5); extended with auto-relationships, scopes, tag linking, embeddings, archiving
ElevenLabs Custom TTS Loryan Strant (@loryanstrant) MIT Voice profile system — patched with v3 mood modulation: stability slider + audio tag prefix injection for non-agent TTS. HACS auto-updates disabled.
microWakeWord Kevin Ahrendt (@kahrendt) — On-device wake word detection for ESP32-S3; custom wake word models trained with this
calendar_utils — — Google Calendar UID access and event deletion for voice CRUD operations

Patterns & Ideas Referenced

Project Author What We Used
Home Generative Agent Lindo St. Angel (@goruck) Context summarization, semantic memory search, multi-model cost tiers, tool error handling, critical action PIN flow (architectural reference — no code copied)
RC Home Assistant Low-VRAM RoyalCities Visible memory via todo list pattern, follow-up conversation state machine
HA Music Voice Control SpotifyPlus brix29 Voice control patterns for Music Assistant integration
Music follow-me concept Phil Hawthorne Presence-based audio routing inspiration
EL-HARP activity prediction Academic paper Feature taxonomy for activity prediction (time bucketing, dwell times, transition probabilities) applied to Markov presence prediction
arc42 architecture template Gernot Starke, Peter Hruschka Document structure inspiration for voice context architecture doc
PEveleigh dispatcher pattern peveleigh Basic conversation.process routing between agent instances; extended into 7-level pipeline-aware dispatcher

Third-Party Blueprints Used

Author Blueprints
Blackshome Sensor light, motion-activated lighting
luuquangvu Memory tool blueprints (local + full LLM)
TheFes Music Assistant LLM voice script (adapted as llm_voice_script.yaml)
Music Assistant Official MA blueprints
SpotifyPlus SpotifyPlus integration blueprints

HA Core Components

Component Role
HA Voice Assistants (Assist Pipelines) Centralized pipeline configuration — routing layer for all 22 persona variants
ElevenLabs (HACS custom + official) Primary TTS via custom component (voice profiles + mood modulation); official integration as secondary entity source
Google Calendar L3 data source for calendar promotion, alarm scheduling, and voice CRUD
IMAP L3 data source for email priority filtering
Aqara (via ha_aqara_devices) FP2 presence detection across 8 zones

Additional Tools & Libraries

Tool Role
OpenRouter LLM API gateway — dynamic per-model pricing, cost tracking
Serper Web search API for agents
sqlite-vec Semantic vector embeddings for L2 memory search
FluidSynth / midiutil Local MIDI synthesis for budget-aware music composition
ha_text_ai Tool-free LLM text generation for non-conversational tasks

Cultural References

  • Young Frankenstein (1974, Mel Brooks) — Project name ("It's pronounced Fronkensteen")
  • Star Trek: Deep Space Nine (Paramount) — Quark persona, Rules of Acquisition naming, Ferengi philosophy throughout
  • Rick and Morty (Adult Swim) — Rick Sanchez persona
  • Deadpool (Marvel / 20th Century Studios) — Deadpool persona
  • Seinfeld (NBC) — Cosmo Kramer persona
  • Doctor Portuondo (Filmin series / Carlo Padial novel) — Doctor Portuondo persona

Tools

  • Claude (Anthropic) — All code generation, architecture design, audit, and documentation
  • Google Gemini — All header image generation (Rick & Morty / Star Trek animation styles)

Thanks

To Jessica — for her patience with this obsession, for coexisting with four AI personalities she never auditioned for, and for being living proof that the best things in a smart home have nothing to do with the technology.


Custom Wake Word Models

Trained models available at mmadalone/microwakeword:

  • Hey Rick / Yo Rick
  • Hey Quark / Yo Quark

But currently using "Okay Nabu" cause mine need better training.

License

This is a personal home automation configuration shared for reference and inspiration. Use whatever's useful to you. If you build on any of the blueprints, a mention would be appreciated but isn't required.

Rule of Acquisition #286: when Morn leaves, it's all over.

About

It's pronounced "Fronkensteen". Five AI personas run your house from a RPi 5

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages