A desktop chat client for local Ollama and llama.cpp models. Built with Electron.
Ported from the OllamaBro browser extension — all the same features, now as a standalone desktop app.
- Fixed llama.cpp responses cutting off mid-stream — normal chat now uses phased connection, first-token, inactivity, and max-stream timeouts instead of a hard 2-minute cap
- Clear cutoff messages for llama.cpp — token-limit and timeout interruptions now show a visible note instead of silently saving a broken answer
- llama.cpp Thinking toggle support — the existing Thinking checkbox now passes
enable_thinkingto compatible GGUF chat templates
- Ollama — chat with any locally installed Ollama model
- llama.cpp — run GGUF models directly via
llama-server; configure binary path, models directory, GPU layers, context size, and server port llama.cppscans GGUFs into a manifest with per-model runtime profiles, inferred capabilities, and automaticmmprojpairing for multimodal setupsllama.cppnow supports the same app-side web search, deep research, memory injection, and explicit memory-save automation used by the Ollama backendllama.cppsession state is persisted so the app can optionally recover the last active GGUF runtime and reuse in-flight loads when startup keep-alive is enabledllama.cppchat streams use phased timeout handling with visible timeout and token-limit notes, plus Thinking toggle support for compatible reasoning templates- Document attachments work with
llama.cppeven when image attachments are unavailable - Model switcher, model management, and dashboard views with live availability checking plus separate local, cloud, and
llama.cppsections - Dot-safe model history storage with legacy migration so dashboard and usage stats stay accurate for future model names
- Kokoro TTS via the local proxy with on-demand model caching, sentence streaming playback, and a built-in self-test endpoint
- Startup diagnostics for Ollama, llama.cpp, memory, and voice prerequisites with guided recovery actions
- Auto-detect context window size per model
- Override context limit manually
- Adjust model parameters per model: temperature, top-p, top-k, repeat penalty, max tokens, seed
- Pull new models from the Ollama registry
- Pull input accepts either a bare model name or pasted
ollama pull .../ollama run ...commands - Update individual models or bulk-update all at once
- Hardware-based model recommendations via llmfit integration
- Full streaming responses with stop button
- Non-destructive regenerate with per-message response version history
- Fork conversations from any message, including into a different model
- Relaunching the app focuses the existing window instead of opening a duplicate app instance
- Thinking/reasoning model support — collapsible reasoning blocks (DeepSeek, QwQ, etc.)
- Multiple conversations per model with search and tag filtering
- Message history navigation with
↑/↓ - Pin messages to preserve them in context
- Per-message actions: copy, read aloud, download as TXT or Markdown, remove
- GitHub Releases update checks with optional startup notifications, a manual check button in Settings, and a compact update popup with version/date info plus
Update now/Restart to installactions for packaged Windows builds - Drag & drop file attachments
- Supported file types: images, PDF, TXT, Markdown, Python, JS, TS, JSON, HTML, CSS, SQL, Shell, YAML, XML, CSV, logs
- Document attachments are chunked and summarized automatically, with only the most relevant excerpts injected into context on each turn
- Scanned PDFs and plain image attachments are OCR-processed automatically, so text-only models can use screenshots, photos, and image-based documents too
- Attached documents stay useful on follow-up questions instead of being pasted in full every time
- Export full conversations as Markdown
- Context meter showing token breakdown (system prompt / search / conversation)
- Top-bar
Chat/Agentworkflow switcher — quickly move between normal conversation and the coding/automation workflow - Web search — augment responses with live Tavily search results
- Deep research — multi-step Exa research pipeline before answering
- Agent mode — autonomous tool-use loop with step-by-step visualization, live in-progress status feedback between tool phases, configurable max steps and permissions, and support for web search, deep research, memory injection, and skill hints in a single run
- Agent capabilities strip — in Agent workflow, pick research mode (
Off / Web / Deep / Auto), memory mode (Off / Inject / Inject + Save), and skills mode (Auto / Manual) - Coding-oriented agent tools — targeted file range reads, codebase search, globbing, in-file replace, single-file patch application, and file system helpers for repo work
- Per-run tool cache — search, glob, and directory results are cached per run and auto-invalidated on writes to avoid redundant rescans
- Agent run history — sidebar shows past and active runs with status indicators and click-to-replay when in Agent mode
- Plan-first approvals — risky shell and file-changing actions now require a plan approval step before execution
- Durable agent run APIs — persisted run metadata plus event log streaming, cancel, and resume endpoints for longer-lived agent workflows
- Renderer run recovery — the desktop UI can reconnect to active runs, replay recent runs, and resume from max-step limits through the durable run API
- Workspace-aware run panels — Agent workflow shows the active workspace, a dedicated changed-files panel, diff previews when patches are available, shell output separated from chat text, and a per-run final summary
- Collapsible dev-tool run panels — each panel (Answer, Timeline, Files, Diffs, Shell) is a foldable section with a count badge and tone-coded border; everything but the Answer starts collapsed so warnings, statuses, and tool calls don't bury the actual response
- Optional YOLO mode — Agent workflow can skip permission and plan prompts for a run while still keeping blocked-path and workspace safety boundaries in place
- Persistent agent progress strip — active runs keep the current step and status visible at the bottom of the response so long multi-step tasks are easier to monitor
- Memory — semantic memory using
nomic-embed-textembeddings; auto-inject and auto-extract toggles; full memory manager with search, add, and clear - Semantic memory search in the memory manager, with provenance like source type, extraction mode, and linked conversation/message metadata
- Auto-extracted memories save directly with provenance and deduplication, and replies can show which memories were used in context
- Voice input — speech-to-text via Whisper (requires Python + faster-whisper)
- Text-to-speech — read responses aloud; choose between Browser (Web Speech API) or Kokoro engine; configurable voice selection
- System prompt editor with token counter
- Persona presets — save, edit, and switch between named system prompts
- Slash commands — custom prompt shortcuts with autocomplete (
/command)
- Dark: Default, Dracula, Tokyo Night, Catppuccin Mocha, Kanagawa, Rosé Pine, Nord, Night Owl, One Dark Pro
- Light: GitHub Light, Solarized Light, Catppuccin Latte, Kanagawa Lotus, Rosé Pine Dawn, Nord Light, Night Owl Light, One Light
| Action | Shortcut |
|---|---|
| Send message | Enter / Ctrl+Enter |
| New chat | Ctrl+N |
| Delete conversation | Ctrl+D |
| Read last response aloud | Ctrl+R |
| Abort generation | Esc |
| Browse message history | ↑ / ↓ |
| Show shortcuts panel | Ctrl+H |
| Toggle voice input | Alt+V |
| Add file | Ctrl+I |
| Toggle web search | Alt+W |
| Toggle deep research | Alt+R |
| Toggle Agent workflow | Alt+A |
| Toggle memory | Ctrl+M |
Download the installer
- Grab the latest Windows installer from GitHub Releases
Prerequisites
Run from source
git clone https://github.com/BorisHrzenjak/ollama_brah.git
cd ollama_brah
npm install
npx electron-rebuild
npm start
electron-rebuildis required becausebetter-sqlite3is a native addon that must be compiled against Electron's bundled Node version. Skipping this step will cause a crash on startup.
Run tests
npm testBuild a distributable
npm run buildProduces a Windows installer in /dist.
Publish a release
git tag -a v1.5.2 -m "Your version notes here"
git push origin v1.5.2The GitHub Actions release workflow builds the installer and uploads it to the matching GitHub Release automatically.
- Ollama must be running before launching the app (
ollama serve) - If Ollama is running on a custom address, set the Ollama server URL in
Settings > Ollamaor setOLLAMA_API_BASE_URLin.env - Pull at least one model first:
ollama pull <model-name> - The internal proxy runs on
localhost:3456— make sure that port is free - Memory feature requires the
nomic-embed-textmodel:ollama pull nomic-embed-text - Voice input requires:
pip install faster-whisper - First OCR use may take longer while English OCR assets are cached locally