About Wallie
Overview
Wallie is an open-source AI streamer that watches and hears your screen and reacts live as any personality you design. It targets faceless content, autonomous AI streams, and real-time commentary, and it runs entirely on your machine with your own API keys.
Built for long real streams instead of short demos, Wallie develops thoughts across minutes, remembers what it covered an hour ago, holds opinions, and drifts between topics like a real conversation. It stays quiet when there is nothing worth saying.
Key Benefits
- Design full personas with identity, voice, humor style, catchphrases, opinions, and taboo topics, and switch between saved profiles from a dropdown.
- Hold hours-long sessions without decay using a rolling summarizer, session notes, cross-session memory, and a paraphrase-aware dedupe engine.
- React to the screen in first person with a SKIP escape hatch and an attention engine that decides which changes deserve deep reactions, glances, or silence.
- Hear what is playing: capture system audio, transcribe speech and lyrics locally, and read music mood and production quality through DSP.
- Play Minecraft autonomously in v2.0, gathering, crafting, building, and fighting with commentary grounded in real game state.
- Drive a VTube Studio Live2D avatar with viseme lip sync, blinks, body motion, and 11 keyword-triggered expression slots.
How It Works
Wallie captures your screen (mss + pHash) and system audio, then fuses chat, vision, and hearing intents into a single prompt through one orchestrator and one conversation history. The LLM streams tokens into a sentence streamer, the TTS pipeline renders PCM16 audio, and the audio player routes it to speakers, OBS, or a virtual cable while the Mood Engine feeds the avatar.
Everything is configured in a browser dashboard at http://127.0.0.1:8765. On Windows you double-click start.bat; on macOS or Linux you run ./start.sh. The first run installs dependencies, then you paste an API key, pick a model, and hit Start.
Use Cases
- Faceless content creators who want an AI personality to commentate gameplay or browsing without showing a face.
- Streamers who want autonomous live commentary that reacts to chat, screen, and audio while they play.
- VTube Studio users who want a Live2D avatar with emotional, mood-reactive animation driven by AI.
- Solo developers and hobbyists who want a fully local, offline streamer using Ollama plus Piper with no API costs.
- Operators running 24/7 autonomous streams on Twitch, YouTube, or Kick.
- Gameplay creators who want unscripted Minecraft playthroughs narrated in character.
Why Choose This Product
Most AI streamer projects fall apart after ten minutes with repetition, short memory, robotic narration, and question loops. Wallie ships fixes for each of these problems through one unified pipeline with a single conversation history and no competing buffers.
It is open source, runs locally on your own keys, and lets you swap every provider without changing code โ from a fully offline setup with Ollama and Piper up to a premium Claude and ElevenLabs stack.
Wallie Pros & Cons
- Open source and runs fully locally on your own API keys
- One-click setup: double-click start.bat, dashboard at 127.0.0.1:8765
- Bring your own everything: 6 LLMs, 3 TTS engines, 3 chat platforms
- Rolling summarizer, session memory, and dedupe prevent long-stream decay
- Play mode autonomously plays Minecraft with grounded commentary
- Screen reactions require a vision-capable LLM model
- Pure-hearing mode needs two extra deps: soundcard and faster-whisper
- Piper TTS voices require a manual one-time download command
- Dashboard should not be exposed to a public network without a reverse proxy
Key Features
Persona design system
Design characters with identity, voice energy, humor style, catchphrases, opinions, and taboo topics, and save multiple personas switchable from a dropdown.
Long-session memory
A rolling summarizer compresses older turns into bullet notes every ~14 segments, while session notes, cross-session memory, and a dedupe engine prevent repetition and decay.
First-person screen vision
Vision uses first-person ownership, a SKIP escape hatch, activity adaptation, and an attention engine that assigns deep reactions, glances, tangents, ignores, and silence beats.
System audio hearing
WASAPI loopback capture with local faster-whisper transcription and pure-numpy DSP that reads key, tempo, instrumentation texture, and production quality of music.
Autonomous game play
Play mode lets Wallie play Minecraft survival on its own using a planning brain, Baritone pathfinding, and a custom Fabric mod for crafting, combat, and a human-like camera.
Emotive Live2D avatar
The VTube Studio avatar runs six animation layers including viseme lip sync, natural blinks, body sway, and 11 keyword-triggered expressions over a single WebSocket.
Browser dashboard
Everything is configured in a browser at http://127.0.0.1:8765 with no YAML files, live status panels, and instant monologue, chat, and vision test previews.
Swappable providers
Six LLM providers, three TTS engines, and three chat platforms can be mixed and matched per profile without changing code.
Single-pipeline architecture
One orchestrator, one conversation history, and one output path prevent the repetition and contradiction seen in multi-path prototypes, with intent priority for chat, vision, and monologue.
Local and private
API keys are stored in .env with restricted permissions, the dashboard binds only to 127.0.0.1, and key values are masked in the UI.
Wallie Pricing
Pricing extracted from the product website and may change. Check the source for current details.
Frequently asked questions about Wallie
How do I install Wallie?
Download the ZIP and unzip it (or git clone the repository), then double-click start.bat on Windows. The first run installs everything it needs and opens the dashboard at http://127.0.0.1:8765. On macOS or Linux, run ./start.sh instead.
How much does it cost to run Wallie?
The software itself is open source and runs locally on your own keys, so there is no license fee. Runtime cost depends on the providers you choose: the free path (Gemini 2.5 Flash plus local Piper TTS) costs $0/hour, a cheap setup with Groq and Fish Audio runs about $1.50/hour, and a premium Claude plus ElevenLabs stack runs about $6.50/hour.
Can I run Wallie fully offline?
Yes. Use Ollama as the local LLM and Piper as the local TTS engine; neither requires an API key. The page shows a fully offline stream as a supported combination alongside premium hosted providers.
Do I need a vision-capable LLM for screen reactions?
Yes. In the dashboard, toggle "Vision capable" ON in the Engine section and use a vision model such as claude-sonnet-4-5, gpt-4o, gemini-2.5-pro, or llama-4-scout on Groq. The SKIP escape hatch is always active and needs no configuration.
How do I route Wallie's audio into OBS?
On Windows, install VB-CABLE, set CABLE Input as your default playback device, and add an Audio Input Capture pointing to CABLE Output in OBS. macOS users can use BlackHole the same way, and Linux users can use PipeWire or PulseAudio loopback.
Can Wallie actually play games?
Yes. v2.0 added Play mode for Minecraft: a planning brain sets goals while Baritone handles pathfinding and mining, and a custom Fabric mod handles crafting, combat, and a human-looking camera. The dashboard installs Fabric and all mods with one click and backs up existing ones first.