# Triage > Capture crashes and errors from your machines, decide which are worth fixing, and suggest fixes. # Overview Source: https://triage.timmo.dev/ Triage reads the systemd journal on each of your machines, picks out crashes, failed units, out-of-memory kills and errors, and groups them into issues on one server. A decision model can then tell you which issues are worth fixing, and a language model can suggest how. ```bash # The issues on this machine, most recently seen first triage collect triage issues # Ask a language model how to fix one triage suggest ``` ## How it fits together
Hosts read the logs
The server groups them into issues
Models decide and suggest
Triage has four roles, and any machine can take on any of them: | Role | Runs | Does | | --- | --- | --- | | Collect | `triage collect --follow --upload` | Reads the journal, redacts what it finds and sends it to the server | | Serve | `triage serve` | Stores events, groups them into issues and enforces the daily AI limits | | Decide | `triage work --decide` or `serve --decide` | Asks a decision model which new issues are worth fixing | | Suggest | `triage work --suggest` or `serve --suggest` | Asks a language model how to fix the issues worth fixing | AI only runs when you ask for it. Decide and suggest are off until you turn them on, and stop at their daily limits. ## Plan first A small always-on box can serve while a machine with more power decides and suggests as a worker, or one machine can do it all. Before installing anything, decide which machine takes each role. **[Plan your hardware](/hardware)** What each role needs, and the setups that work best. ## Get started 1. **Install** The Arch packages, the Home Assistant app, the container or a release build. See [Install](/install). 2. **Run a server** Where events and issues live. See [Run a server](/setup/server). 3. **Add hosts** Collect from each machine. See [Add hosts](/setup/hosts). 4. **Add workers** Optional: run the models on another machine. See [Add workers](/setup/workers). ## Reference **[Configuration](/configuration)** Every setting. **[Issues](/issues)** States, labels and how models are measured. **[Choices](/choices)** Hosting and models, local or hosted. **[Privacy](/privacy)** What's captured, what's redacted and where it goes. **[Libraries](/libraries)** Talk to a server from your own Effect app. **[Commands](/commands)** Every command and flag. ## Machine-readable docs | URL | Use | | --- | --- | | https://triage.timmo.dev/llms.txt | Compact page index | | https://triage.timmo.dev/llms-full.txt | Full docs bundle | | https://triage.timmo.dev/mcp | Hosted MCP (search, page and navigation) | | `/{route}.md` | One page as Markdown | --- # Choices Source: https://triage.timmo.dev/choices Triage doesn't tie you to any vendor. Every part can run on your own hardware or on a service you pick, and none of them is the intended one with the rest as fallbacks. Mix them however suits you: the server on one machine, decision models on another, and a hosted language model, or everything on one box with nothing leaving your network. See [Privacy](/privacy) for exactly what each of them is sent, and [Configuration](/configuration#models) for the settings. ## Hosting the server | Option | Where data goes | Cost | | --- | --- | --- | | Arch Linux user service | Your machine | Free | | Container, with Docker Compose | Your machine, plus Cloudflare if you use a [Cloudflare Tunnel](https://developers.cloudflare.com/cloudflare-one/networks/connectors/cloudflare-tunnel/) | Free | | Home Assistant app | Your Home Assistant | Free | | Cloudflare (planned) | Your Cloudflare account | [Cloudflare Workers pricing](https://developers.cloudflare.com/workers/platform/pricing/) | Hosts and workers only need the server's URL, so you can move it later without changing anything else. ## Decision models Decision models decide which issues are worth fixing. Set `TRIAGE_DECISION_PROVIDER` to `typesafe` for any TypeSafe System One API, or `cloudflare` for Clef. | Option | Where data goes | Cost | | --- | --- | --- | | [Ollaya](https://ollaya.dev/library), such as [`laya`](https://ollaya.dev/library/laya) or [`winnow`](https://ollaya.dev/library/winnow) | Your machine | Free | | [Ollama](https://ollama.com/search?c=decision) 0.35 or later, such as [`nimble`](https://ollama.com/library/nimble) or [`tev1`](https://ollama.com/library/tev1) | Your machine | Free | | [TypeSafe](https://typesafe.ai/)'s API | TypeSafe | TypeSafe's pricing | | Jev on [OpenCode Zen](https://opencode.ai/docs/zen/) | OpenCode | [OpenCode Zen's pricing](https://opencode.ai/docs/zen/#pricing) | | Clef on Cloudflare Workers AI, [`clef`](https://developers.cloudflare.com/workers-ai/models/clef/) or [`clef-flash`](https://developers.cloudflare.com/workers-ai/models/clef-flash/) | Your Cloudflare account | [Workers AI's](https://developers.cloudflare.com/workers-ai/platform/pricing/) free 10,000 Neurons a day, then its pricing | ## Language models Language models suggest fixes. Set `TRIAGE_LLM_PROVIDER` to `openai` for any OpenAI-compatible API, `anthropic` for any Anthropic-compatible one, or `cloudflare` for Workers AI. | Option | Where data goes | Cost | | --- | --- | --- | | [Ollama](https://ollama.com/library), [LM Studio](https://lmstudio.ai/models) or [llama.cpp](https://github.com/ggml-org/llama.cpp) | Your machine | Free | | [OpenAI](https://developers.openai.com/api/docs/models) | OpenAI | [OpenAI's pricing](https://openai.com/business/pricing/#api) | | [Anthropic](https://platform.claude.com/docs/en/models/overview) | Anthropic | [Anthropic's pricing](https://platform.claude.com/docs/en/about-claude/pricing) | | [OpenRouter](https://openrouter.ai/models) | OpenRouter and the model's provider | Each model's price on OpenRouter | | [OpenCode Zen](https://opencode.ai/docs/zen/), through either API | OpenCode | [OpenCode Zen's pricing](https://opencode.ai/docs/zen/#pricing) | | [GitHub Models](https://docs.github.com/en/github-models) | GitHub | GitHub's free allowance, then its pricing | | [Workers AI](https://developers.cloudflare.com/workers-ai/models/) | Your Cloudflare account | [Workers AI's](https://developers.cloudflare.com/workers-ai/platform/pricing/) free 10,000 Neurons a day, then its pricing | One Ollama can serve both the decision model and the language model. Planned: - A coding agent such as OpenCode, Pi, Cursor, Claude Code, Codex, Copilot or Gemini, which reaches whatever providers it's set up with, such as a Copilot subscription you already have. - Home Assistant's AI Task action, which uses whichever AI provider your Home Assistant is set up with. ## Limits Decide and suggest are off until you turn them on, and stop at their daily limits (`TRIAGE_DECIDE_DAILY` and `TRIAGE_SUGGEST_DAILY`), shared by the server and all its workers. Each suggestion is capped at 4,096 tokens, so its cost stays predictable on any paid API. --- # Commands Source: https://triage.timmo.dev/commands Each command has its own page with its help, as `triage --help` prints it. | Command | Alias | | --- | --- | | [`collect`](/commands/collect) | None | | [`issues`](/commands/issues) | None | | [`upload`](/commands/upload) | None | | [`serve`](/commands/serve) | None | | [`work`](/commands/work) | None | | [`hosts`](/commands/hosts) | None | | [`admins`](/commands/admins) | None | | [`workers`](/commands/workers) | None | | [`decide`](/commands/decide) | None | | [`suggest`](/commands/suggest) | None | | [`label`](/commands/label) | None | | [`resolve`](/commands/resolve) | None | | [`mute`](/commands/mute) | None | | [`reopen`](/commands/reopen) | None | | [`agreement`](/commands/agreement) | None | ## Global flags ```text DESCRIPTION Capture crashes and errors from your machines, decide which are worth fixing, and suggest fixes USAGE triage [flags] GLOBAL FLAGS --help, -h Show help information --version, -v Show version information --wizard Start wizard mode for a command --completions Print shell completion script (choices: bash, zsh, fish, sh) --log-level Sets the minimum log level (choices: all, trace, debug, info, warn, warning, error, fatal, none) SUBCOMMANDS collect Collect crashes, failures and errors from this machine's journal since the last run issues List issues, most recently seen first upload Send collected events to the server at $TRIAGE_SERVER, authenticating with $TRIAGE_TOKEN serve Run the triage server over HTTP, which collects events from enrolled hosts. Use a reverse proxy or Cloudflare for HTTPS work Decide on issues and suggest fixes for the server at $TRIAGE_SERVER, authenticating with $TRIAGE_WORKER_TOKEN, within the server's daily limits. Needs --decide, --suggest or both hosts Manage the hosts that can send events admins Manage the admins that can read issues and manage tokens workers Manage the workers that can decide on issues and suggest fixes for this server decide Ask a decision model whether the server's new issues are worth fixing, storing the answers without acting on them suggest Ask a language model how to fix some of the server's issues, from their redacted events only, and store its suggestions label Label one of the server's issues by hand, to measure decision models against resolve Resolve issues once they're fixed. One that happens again opens as regressed, and is decided on again mute Mute issues, so they're never decided on or suggested fixes for, however often they happen reopen Reopen resolved or muted issues agreement Compare each decision model with the hand labels: how often it's sure enough to act on, and how often it's right when it is ``` --- # triage admins Source: https://triage.timmo.dev/commands/admins Every `triage admins` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage admins` ```text DESCRIPTION Manage the admins that can read issues and manage tokens USAGE triage admins [flags] ``` ## `triage admins add` ```text DESCRIPTION Add an admin and print its token USAGE triage admins add [flags] ARGUMENTS name string A name for the admin, such as aidan FLAGS --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` ## `triage admins list` ```text DESCRIPTION List each an admin, oldest first USAGE triage admins list [flags] FLAGS --json Print JSON --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` ## `triage admins remove` ```text DESCRIPTION Remove an admin, revoking its token USAGE triage admins remove [flags] ARGUMENTS name string A name for the admin, such as aidan FLAGS --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` --- # triage agreement Source: https://triage.timmo.dev/commands/agreement Every `triage agreement` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage agreement` ```text DESCRIPTION Compare each decision model with the hand labels: how often it's sure enough to act on, and how often it's right when it is USAGE triage agreement [flags] FLAGS --json Print JSON ``` --- # triage collect Source: https://triage.timmo.dev/commands/collect Every `triage collect` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage collect` ```text DESCRIPTION Collect crashes, failures and errors from this machine's journal since the last run USAGE triage collect [flags] FLAGS --follow, -f Keep collecting new entries as they're written --upload, -u Send new events to $TRIAGE_SERVER after each batch, keeping them to retry when it can't be reached --json Print JSON ``` --- # triage decide Source: https://triage.timmo.dev/commands/decide Every `triage decide` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage decide` ```text DESCRIPTION Ask a decision model whether the server's new issues are worth fixing, storing the answers without acting on them USAGE triage decide [flags] FLAGS --provider choice Decide through any TypeSafe System One API, such as Ollaya or Ollama locally, or with Clef on Cloudflare using $CLOUDFLARE_ACCOUNT_ID and $CLOUDFLARE_API_TOKEN (choices: typesafe, cloudflare) --url string The System One API: Ollaya by default, Ollama at http://127.0.0.1:11434/v1, or a hosted one with $TRIAGE_DECISION_API_KEY --model, -m string The decision model: laya by default through System One, clef-flash with Cloudflare --limit, -n integer The most issues to decide on --json Print JSON ``` --- # triage hosts Source: https://triage.timmo.dev/commands/hosts Every `triage hosts` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage hosts` ```text DESCRIPTION Manage the hosts that can send events USAGE triage hosts [flags] ``` ## `triage hosts add` ```text DESCRIPTION Add a host and print its token USAGE triage hosts add [flags] ARGUMENTS name string A name for the host that doesn't identify the machine, such as desktop FLAGS --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` ## `triage hosts list` ```text DESCRIPTION List each a host, oldest first USAGE triage hosts list [flags] FLAGS --json Print JSON --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` ## `triage hosts remove` ```text DESCRIPTION Remove a host, revoking its token USAGE triage hosts remove [flags] ARGUMENTS name string A name for the host that doesn't identify the machine, such as desktop FLAGS --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` --- # triage issues Source: https://triage.timmo.dev/commands/issues Every `triage issues` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage issues` ```text DESCRIPTION List issues, most recently seen first USAGE triage issues [flags] FLAGS --limit, -n integer The most issues to show --server List the server's issues from $TRIAGE_SERVER_DB, for decide, label and suggest --json Print JSON ``` --- # triage label Source: https://triage.timmo.dev/commands/label Every `triage label` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage label` ```text DESCRIPTION Label one of the server's issues by hand, to measure decision models against USAGE triage label [flags] ARGUMENTS issue string The issue's ID verdict choice worth: a real fault worth fixing; noise: expected, harmless or caused by the user ``` --- # triage mute Source: https://triage.timmo.dev/commands/mute Every `triage mute` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage mute` ```text DESCRIPTION Mute issues, so they're never decided on or suggested fixes for, however often they happen USAGE triage mute [flags] ARGUMENTS issue... string The IDs of the issues FLAGS --server string Change the issues on the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` --- # triage reopen Source: https://triage.timmo.dev/commands/reopen Every `triage reopen` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage reopen` ```text DESCRIPTION Reopen resolved or muted issues USAGE triage reopen [flags] ARGUMENTS issue... string The IDs of the issues FLAGS --server string Change the issues on the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` --- # triage resolve Source: https://triage.timmo.dev/commands/resolve Every `triage resolve` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage resolve` ```text DESCRIPTION Resolve issues once they're fixed. One that happens again opens as regressed, and is decided on again USAGE triage resolve [flags] ARGUMENTS issue... string The IDs of the issues FLAGS --server string Change the issues on the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` --- # triage serve Source: https://triage.timmo.dev/commands/serve Every `triage serve` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage serve` ```text DESCRIPTION Run the triage server over HTTP, which collects events from enrolled hosts. Use a reverse proxy or Cloudflare for HTTPS USAGE triage serve [flags] FLAGS --hostname string The address to listen on --port, -p integer The port to listen on --trust-proxy Trust X-Forwarded-Host and X-Forwarded-For from a reverse proxy; only when the proxy is the sole way in --ingress-port integer A second port for Home Assistant ingress, which only answers --ingress-from and needs no admin token --ingress-from string The only address the ingress port answers, Home Assistant's Supervisor --decide-daily integer The most issues each decision model may decide on in any 24 hours, here with --decide or by workers --suggest-daily integer The most suggestions each language model may make in any 24 hours, here with --suggest or by workers --decide Ask a decision model about new issues every few minutes, keeping the answers without acting on them. Off unless set --provider choice Decide through any TypeSafe System One API, such as Ollaya or Ollama locally, or with Clef on Cloudflare using $CLOUDFLARE_ACCOUNT_ID and $CLOUDFLARE_API_TOKEN (choices: typesafe, cloudflare) --url string The System One API: Ollaya by default, Ollama at http://127.0.0.1:11434/v1, or a hosted one with $TRIAGE_DECISION_API_KEY --model, -m string The decision model: laya by default through System One, clef-flash with Cloudflare --suggest Ask a language model every 15 minutes how to fix issues the decision model clearly rates worth fixing, keeping its suggestions. Off unless set --llm-provider choice Any OpenAI-compatible or Anthropic-compatible API at --llm-url, or Workers AI with $CLOUDFLARE_ACCOUNT_ID and $CLOUDFLARE_API_TOKEN (choices: openai, anthropic, cloudflare) --llm-url string The API, such as https://openrouter.ai/api/v1 or https://opencode.ai/zen/v1, with $TRIAGE_LLM_API_KEY when it needs one. Defaults to OpenAI's or Anthropic's own --llm-model string The language model, such as @cf/zai-org/glm-4.7-flash on Workers AI ``` --- # triage suggest Source: https://triage.timmo.dev/commands/suggest Every `triage suggest` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage suggest` ```text DESCRIPTION Ask a language model how to fix some of the server's issues, from their redacted events only, and store its suggestions USAGE triage suggest [flags] ARGUMENTS issue... string The IDs of the issues to suggest fixes for FLAGS --provider choice Any OpenAI-compatible or Anthropic-compatible API at --url, or Workers AI with $CLOUDFLARE_ACCOUNT_ID and $CLOUDFLARE_API_TOKEN (choices: openai, anthropic, cloudflare) --url string The API, such as https://openrouter.ai/api/v1 or https://opencode.ai/zen/v1, with $TRIAGE_LLM_API_KEY when it needs one. Defaults to OpenAI's or Anthropic's own --model, -m string The language model, such as @cf/zai-org/glm-4.7-flash on Workers AI --json Print JSON ``` --- # triage upload Source: https://triage.timmo.dev/commands/upload Every `triage upload` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage upload` ```text DESCRIPTION Send collected events to the server at $TRIAGE_SERVER, authenticating with $TRIAGE_TOKEN USAGE triage upload [flags] FLAGS --json Print JSON ``` --- # triage work Source: https://triage.timmo.dev/commands/work Every `triage work` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage work` ```text DESCRIPTION Decide on issues and suggest fixes for the server at $TRIAGE_SERVER, authenticating with $TRIAGE_WORKER_TOKEN, within the server's daily limits. Needs --decide, --suggest or both USAGE triage work [flags] FLAGS --decide Ask a decision model about new issues every few minutes, keeping the answers without acting on them. Off unless set --provider choice Decide through any TypeSafe System One API, such as Ollaya or Ollama locally, or with Clef on Cloudflare using $CLOUDFLARE_ACCOUNT_ID and $CLOUDFLARE_API_TOKEN (choices: typesafe, cloudflare) --url string The System One API: Ollaya by default, Ollama at http://127.0.0.1:11434/v1, or a hosted one with $TRIAGE_DECISION_API_KEY --model, -m string The decision model: laya by default through System One, clef-flash with Cloudflare --suggest Ask a language model every 15 minutes how to fix issues the decision model clearly rates worth fixing, keeping its suggestions. Off unless set --llm-provider choice Any OpenAI-compatible or Anthropic-compatible API at --llm-url, or Workers AI with $CLOUDFLARE_ACCOUNT_ID and $CLOUDFLARE_API_TOKEN (choices: openai, anthropic, cloudflare) --llm-url string The API, such as https://openrouter.ai/api/v1 or https://opencode.ai/zen/v1, with $TRIAGE_LLM_API_KEY when it needs one. Defaults to OpenAI's or Anthropic's own --llm-model string The language model, such as @cf/zai-org/glm-4.7-flash on Workers AI ``` --- # triage workers Source: https://triage.timmo.dev/commands/workers Every `triage workers` command and its help, as `--help` prints it. Each also accepts the [global flags](/commands#global-flags). ## `triage workers` ```text DESCRIPTION Manage the workers that can decide on issues and suggest fixes for this server USAGE triage workers [flags] ``` ## `triage workers add` ```text DESCRIPTION Add a worker and print its token USAGE triage workers add [flags] ARGUMENTS name string A name for the worker that doesn't identify the machine, such as desktop FLAGS --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` ## `triage workers list` ```text DESCRIPTION List each a worker, oldest first USAGE triage workers list [flags] FLAGS --json Print JSON --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` ## `triage workers remove` ```text DESCRIPTION Remove a worker, revoking its token USAGE triage workers remove [flags] ARGUMENTS name string A name for the worker that doesn't identify the machine, such as desktop FLAGS --server string Manage the server at this URL as the admin in $TRIAGE_ADMIN_TOKEN, instead of the server database on this machine ``` --- # Configuration Source: https://triage.timmo.dev/configuration Every setting comes from an environment variable. The user services read theirs from `~/.config/triage/agent.env`, `server.env` and `worker.env`. ## Options file Any setting can also come from a JSON file at `TRIAGE_OPTIONS`, with lowercase keys and no `TRIAGE_` prefix: ```json { "decide": true, "decide_daily": 10, "cloudflare_api_token": "..." } ``` Environment variables win over the file. The Home Assistant app uses this for its options. ## Hosts | Variable | Default | | | --- | --- | --- | | `TRIAGE_SERVER` | none | The server's URL | | `TRIAGE_TOKEN` | none | The host's token, from `triage hosts add` | | `TRIAGE_REDACT_NAMES` | none | Other names to redact, comma-separated, such as a GitHub account | | `TRIAGE_DB` | `$XDG_STATE_HOME/triage/triage.db` | The host's own store, which also holds events until they're sent | ## Server | Variable | Default | | | --- | --- | --- | | `TRIAGE_HOSTNAME` | `127.0.0.1` (`::` in the container, IPv4 and IPv6) | Address to listen on | | `TRIAGE_PORT` | `7171` | Port to listen on | | `TRIAGE_SERVER_DB` | `$XDG_STATE_HOME/triage/server.db` (`/data/server.db` in the container) | Server database | | `TRIAGE_TRUST_PROXY` | `false` | Trust `X-Forwarded-Host` and `X-Forwarded-For`. Only turn this on when the proxy is the only way to reach the server | | `TRIAGE_SERVER_ADMIN_TOKEN` | none | An admin token the server always accepts, at least 32 characters, for managing tokens with `--server`. It isn't stored or listed | | `TRIAGE_INGRESS_PORT` | none (`8099` in the Home Assistant app) | A second port for Home Assistant ingress, which treats every request as an admin's and only answers `TRIAGE_INGRESS_FROM` | | `TRIAGE_INGRESS_FROM` | `172.30.32.2`, Home Assistant's Supervisor | The only address the ingress port answers | | `TRIAGE_DECIDE_DAILY` | `20` | The most issues each decision model may decide on in any 24 hours, by the server and its workers together | | `TRIAGE_SUGGEST_DAILY` | `5` | The most suggestions each language model may make in any 24 hours, by the server and its workers together | ## Workers and admins | Variable | Default | | | --- | --- | --- | | `TRIAGE_SERVER` | none | The server's URL | | `TRIAGE_WORKER_TOKEN` | none | The worker's token, from `triage workers add` | | `TRIAGE_ADMIN_TOKEN` | none | An admin token, for `--server` on the token and issue commands | ## Models These apply wherever the models run: `triage work`, `triage serve`, and the one-off `decide` and `suggest` commands. | Variable | Default | | | --- | --- | --- | | `TRIAGE_DECIDE` | `false` | Ask a decision model about new issues every 5 minutes. The answers are only stored for now, to compare models | | `TRIAGE_DECISION_PROVIDER` | `typesafe` | `typesafe` for any TypeSafe System One API, such as Ollaya or Ollama 0.35+, or `cloudflare` for Clef on Workers AI | | `TRIAGE_DECISION_URL` | `http://127.0.0.1:11435/v1` (Ollaya) | The System One API, such as `http://127.0.0.1:11434/v1` for Ollama | | `TRIAGE_DECISION_API_KEY` | none | Key for a hosted System One API, such as TypeSafe or OpenCode Zen | | `TRIAGE_DECISION_MODEL` | `laya` through System One, `clef-flash` with Cloudflare | Decision model, such as `laya` or `winnow` on [Ollaya](https://ollaya.dev/library), or `nimble` on [Ollama](https://ollama.com/search?c=decision) | | `TRIAGE_SUGGEST` | `false` | Suggest fixes every 15 minutes for issues the decision model rates worth fixing with at least 0.8 probability. Needs `TRIAGE_LLM_MODEL` | | `TRIAGE_LLM_PROVIDER` | `openai` | `openai` for any OpenAI-compatible API, `anthropic` for any Anthropic-compatible one, or `cloudflare` for Workers AI | | `TRIAGE_LLM_URL` | OpenAI's or Anthropic's own | The API, such as `https://openrouter.ai/api/v1`, `https://opencode.ai/zen/v1` or `http://127.0.0.1:11434/v1` for Ollama | | `TRIAGE_LLM_API_KEY` | none | Key for the API, when it needs one | | `TRIAGE_LLM_MODEL` | none | Language model for fix suggestions, such as `@cf/google/gemma-4-26b-a4b-it` on Workers AI | | `CLOUDFLARE_ACCOUNT_ID` | none | Your Cloudflare account, for either provider set to `cloudflare` | | `CLOUDFLARE_API_TOKEN` | none | A token with Workers AI access, for either provider set to `cloudflare` | On the `serve` and `work` command lines, the language model settings are `--llm-provider`, `--llm-url` and `--llm-model`, since `--provider`, `--url` and `--model` are the decision model's. --- # Plan your hardware Source: https://triage.timmo.dev/hardware Triage has four parts, and each one can run on whichever machine suits it. They need very different amounts of hardware, so it's worth deciding where each goes before you set anything up. ## At a glance | Part | Where it runs | Needs | Best on | | --- | --- | --- | --- | | Collect | Every Linux machine you want to watch | systemd's journal, and very little CPU or memory | Each machine itself | | Serve | One machine every host can reach | Little CPU or memory, a small SQLite database, and to stay on | A small always-on box, such as a Home Assistant Green, a home server or a VPS | | Decide | The server or a worker | A decision model, on whatever's available: a GPU, a CPU or a hosted service | A GPU if you have one, otherwise a CPU or a hosted service | | Suggest | The server or a worker | A language model, on a GPU with plenty of memory or a hosted service | A GPU with plenty of memory if you have one, otherwise a hosted service | Decide and suggest are optional, and off until you turn them on. Collect and serve on their own already group everything into issues. ## Collect The agent runs on each host as a user service. It follows the journal, so it needs systemd. Crashes come from systemd-coredump's entries, so they only show up where that's turned on. It uses very little CPU or memory, and keeps events on the host while the server is unreachable, so laptops that come and go are fine. ## Serve The server stores events, groups them into issues and keeps the daily AI limits. It's light: it runs on a Home Assistant Green. What matters more is that it's always on and every host can reach it. Hosts hold on to their events while it's down, but nothing new shows up until it's back. Small boards like the Green or Yellow can serve, but they're rarely up to running models, so leave decide and suggest to a [worker](/setup/workers) on those. ## Decide A decision model reads each new issue and rates whether it's worth fixing. It never writes text, so it's quick and small, and runs on whatever you have: - **A GPU** answers fastest and fits the more accurate models. - **A CPU** is slower, but small models keep up easily: the daily limit is 20 issues by default. - **A hosted service** needs no hardware at all, at that service's prices. For a sense of scale, here are [Ollaya's own figures](https://ollaya.dev/results) for five questions on two of its models, measured on an NVIDIA RTX 4090 and a 24-core x86 CPU: | Model | GPU | CPU | | --- | --- | --- | | [`laya`](https://ollaya.dev/library/laya) (the default) | About 10 ms | 0.2 to 0.4 s | | [`winnow:e4b`](https://ollaya.dev/library/winnow:e4b) | About 90 ms | About 5 s | Ollaya uses NVIDIA GPUs and Apple silicon, and NVIDIA is what it's tested on most. Small models like `laya` run well on a CPU, but larger ones like `winnow:e4b` really want a GPU. Ollaya is one option among several, local and hosted. Browse [Ollaya's models](https://ollaya.dev/library) and [Ollama's decision models](https://ollama.com/search?c=decision), and see [Choices](/choices#decision-models) for every option. ## Suggest A language model writes a suggested fix for the issues the decision model rates worth fixing. This is the heaviest part, and the one where the model you pick makes the most difference: bigger models give better suggestions. - **A GPU** with enough memory for the model runs one locally. A model that doesn't fit runs partly on the CPU and gets much slower. - **A hosted service** needs no hardware, and offers models too large to run at home, at that service's prices. Suggestions are limited to 5 a day by default and checked for every 15 minutes, so a slow local model still keeps up. More ways to suggest fixes are planned, such as a coding agent you already use or Home Assistant's AI Task; see [Choices](/choices#language-models). ## Putting it together | Setup | Serve | Decide and suggest | | --- | --- | --- | | Small box plus a desktop | A Home Assistant Green, home server or VPS | A desktop with a GPU, as a [worker](/setup/workers). Nothing queues up while it's off | | One machine | A desktop or home server with a GPU | The same machine, with `triage serve --decide --suggest` | | No GPU anywhere | Any always-on machine | A small decision model on the CPU, or hosted, and a hosted language model | Local models keep everything on your network. Hosted ones are only ever sent redacted events; see [Privacy](/privacy) for exactly what, and [Choices](/choices) for every option. --- # Install Source: https://triage.timmo.dev/install Triage is a single Linux binary for x86_64 and aarch64. The server also comes as a container image and a Home Assistant app. ## Arch Linux Packages are published to the unofficial `timmo` pacman repository. Add it before the other repository sections in `/etc/pacman.conf`: ```ini [timmo] SigLevel = PackageRequired DatabaseOptional TrustedOnly Server = https://packages.timmo.dev/$arch ``` Then install one of the two packages: | Package | Built from | | --- | --- | | `triage-bin` | The latest release | | `triage-git` | Every push to `main` | ```bash sudo pacman -Syu triage-bin ``` Both install the `triage` command, shell completions and three user services: `triage-agent.service`, `triage-server.service` and `triage-worker.service`. None is enabled on install, since each machine takes on different roles. On upgrade, the package restarts the ones that are running. ## Home Assistant The Home Assistant app runs the server on Home Assistant OS, for amd64 and aarch64. Add this repository to the app store: [![Add the repository to Home Assistant](https://my.home-assistant.io/badges/supervisor_add_addon_repository.svg)](https://my.home-assistant.io/redirect/supervisor_add_addon_repository/?repository_url=https%3A%2F%2Fgithub.com%2Ftimmo001%2Ftriage) then install Triage, set its admin token and start it. See [Run a server](/setup/server#home-assistant). ## Container `ghcr.io/timmo001/triage` runs the server, for amd64 and arm64. See [Run a server](/setup/server#container). ## Debian, Ubuntu and Fedora Each [GitHub release](https://github.com/timmo001/triage/releases) has `.deb` and `.rpm` packages: ```bash # Debian and Ubuntu sudo apt install ./triage__amd64.deb # Fedora sudo dnf install ./triage--1.x86_64.rpm ``` ## Release archive Each release also has a `triage--linux-.tar.gz` archive with just the binary. Put it somewhere on your `PATH`: ```bash tar -xzf triage--linux-x86_64.tar.gz install -Dm755 triage ~/.local/bin/triage ``` Release assets come with a `SHA256SUMS` file and a Sigstore bundle. To check an asset was built by this repository's release workflow: ```bash gh attestation verify triage--linux-x86_64.tar.gz --repo timmo001/triage ``` ## Build from source You need [mise](https://mise.jdx.dev), which installs the pinned Bun and Node versions: ```bash git clone https://github.com/timmo001/triage.git cd triage mise install mise run build ``` The binary is written to `dist/triage`. ## Next steps - [Run a server](/setup/server). - [Add the machines](/setup/hosts) to collect from. --- # Issues Source: https://triage.timmo.dev/issues Triage groups events into issues by what went wrong, not when or where: - Crashes group by executable and signal, plus their top stack frames when there are any. - Failed units group by unit and how they failed. Transient scopes, `systemd-run` units, session scopes and units with numbered instances each group as one. - Out-of-memory kills group by the process that was killed. - Errors group by the program that logged them and the message, with its numbers, paths, IDs and addresses taken out. The same fault on two hosts is one issue, with events from both. ```bash triage issues # this host's issues triage issues --server # the server's, on the server ``` ## States | State | Means | | --- | --- | | New | First seen in the last week | | Ongoing | Seen before that, and still open | | Regressed | Happened again after it was resolved, in the last week | | Resolved | Fixed, as far as you know | | Muted | Hidden from decision and language models, however often it happens | ```bash triage resolve ... triage mute ... triage reopen ... ``` An event later than an issue's resolution reopens it as regressed, and decision models look at it again. Add `--server ` to change a remote server's issues as the admin in `TRIAGE_ADMIN_TOKEN`. ## In a browser The server has a web page at its own URL, such as `http://localhost:7171/`. It lists the server's issues, filtered by state, and each issue's hosts and latest events, with buttons to resolve, mute, reopen or unmute it. Sign in with an [admin token](/setup/server#tokens). It's kept in that browser until you sign out. In the [Home Assistant app](/setup/server#home-assistant), open **Triage** in Home Assistant's sidebar instead, with no token needed. ## Decisions A decision model answers three questions about each new or regressed issue: whether it's worth fixing, how severe it is and what likely caused it, each with probabilities. The answers are only stored for now, to compare models before they drive anything. ```bash triage decide # decide on the server's new issues now ``` ## Labels and agreement Label issues by hand to measure decision models against your own judgement: ```bash triage label worth triage label noise triage agreement ``` `agreement` shows, for each model, how many labelled issues it was sure about (worth at least 0.8 or at most 0.2), and how often it was right when it was. ## Suggestions ```bash triage suggest ... ``` asks a language model how to fix issues, and stores its answers. It sends the same trimmed, redacted description decision models get, plus the unit and the lines it logged before it failed, and caps each response at 4,096 tokens, thinking included. Each suggestion records the events it was based on. --- # Libraries Source: https://triage.timmo.dev/libraries Triage is built from Effect v4 libraries, published to npm and JSR. They work under Bun and Node. Early development: the API will change before 1.0. | Package | Use it to | | --- | --- | | [`@timmo001/effect-triage-client`](https://github.com/timmo001/triage/tree/main/packages/client) | Talk to a triage server over its HTTP API | | [`@timmo001/effect-triage`](https://github.com/timmo001/triage/tree/main/packages/effect-triage) | Use the event, issue and API schemas every part of triage shares | The CLI uses the same API, so anything it does against a server, your app can do too. ## Client ```bash bun add @timmo001/effect-triage-client effect ``` `effect` is a peer dependency, so install the Effect v4 version your app already uses. `TriageClient` is an Effect service built on `effect/http-api`. `TriageClient.layer({ url, token })` makes a typed client for one server, authenticating with a host, worker or admin token. It retries transient failures a few times. ```ts import { TriageClient } from "@timmo001/effect-triage-client"; import { Console, Effect, Layer, Redacted } from "effect"; import { FetchHttpClient } from "effect/http"; const program = Effect.gen(function* () { const client = yield* TriageClient; const issues = yield* client.issues.list({ query: { limit: 10 } }); for (const issue of issues) { yield* Console.log(`${issue.state} ${issue.count} ${issue.title}`); } }); const client = TriageClient.layer({ url: "https://triage.example.com", token: Redacted.make(process.env.TRIAGE_ADMIN_TOKEN ?? ""), }).pipe(Layer.provide(FetchHttpClient.layer)); program.pipe(Effect.provide(client), Effect.runPromise); ``` ## API | Group | Token | Endpoints | | --- | --- | --- | | `ingest` | Host | Send events | | `issues` | Admin | List issues, get one with its events, and set its state | | `work` | Worker | Fetch issues to decide on or suggest for, and send back answers | | `tokens` | Admin | List, add and remove host, worker and admin tokens | | `system` | None | Health | The server serves its OpenAPI document at `/api/openapi.json`. --- # Privacy Source: https://triage.timmo.dev/privacy Triage reads your system journal, which holds a lot of personal detail. This explains what it keeps, what it removes and where anything goes. There's no triage service or account: everything runs on machines and services you choose, and nothing is sent to the developer. ## What's captured Hosts read the systemd journal and keep only: - crashes (systemd-coredump), with the crashing thread's function and library names - unit failures - out-of-memory kills, from the kernel and systemd-oomd - anything logged at error priority or worse For failures and crashes, hosts also keep the last 10 lines the unit logged before it, to give suggestions something to go on. Each event keeps its message, the program and unit that logged it, its severity and when it happened. Everything else in the journal entry is dropped. ## What's redacted Redaction happens on the host, when an event is captured, before it's stored or sent anywhere. Every text field, including stack frames and the logged lines, goes through it. It replaces: - passwords, tokens, API keys and `Bearer` credentials - email addresses - home directories, as `~` - the machine's hostname and every regular user's name - any other names in `TRIAGE_REDACT_NAMES`, comma-separated, such as a GitHub account that shows up in repository URLs - IPv4 and IPv6 addresses - MAC addresses - UUIDs, and other long hex or token-like strings - serial numbers - Wi-Fi network names The journal cursor and boot ID are replaced with hashes, so events don't carry machine identifiers either. When the server receives events, it sets their host to the name you enrolled the host with, never the machine's own hostname. Redaction is pattern based, so it can miss something unusual. Run `triage collect`, then look through what it stored with `triage issues`, before turning on anything that sends data off the machine. ## Where data goes | What | Where | What it sends | | --- | --- | --- | | Hosts | The triage server you set in `TRIAGE_SERVER` | Redacted events | | Server | Nowhere, unless you turn on decide or suggest | | | Workers | The triage server they're enrolled with | Decisions and suggestions | | Decision models, with `TRIAGE_DECIDE` | The decision model API you choose | A redacted description of the issue | | Language models, with `TRIAGE_SUGGEST` or `triage suggest` | The language model API you choose | The same description, plus the unit and its logged lines | The description an issue sends is its kind, title and event count, up to 3 distinct messages and, for crashes, the top 5 stack frames. Messages and logged lines are cut at 300 characters. It never includes host names, event IDs or timestamps. Nothing is sent to a model unless you turn decide or suggest on, or run `triage decide` or `triage suggest` yourself, and the automatic runs stop at their daily limits. With models on your own network, such as Ollaya or Ollama on a [worker](/setup/workers), nothing leaves your network at all. See [Choices](/choices) for where each option sends data. ## What the server stores The server stores the redacted events, the issues they're grouped into, and any decisions and suggestions. Tokens are stored as hashes, so the database doesn't hold a usable token. `TRIAGE_SERVER_ADMIN_TOKEN` isn't stored at all. --- # Add hosts Source: https://triage.timmo.dev/setup/hosts A host is a machine triage collects from. It reads the journal, keeps crashes, failed units, out-of-memory kills and errors, redacts them and sends them to the server. ## Enrol On the server, or from anywhere with [`--server`](/setup/server#manage-tokens-remotely), add the host by name: ```bash triage hosts add laptop ``` The name is how the host shows up on the server, never its own hostname. Use lowercase letters, numbers and hyphens. ## Run the agent On the host, put the server's URL and the printed token in `~/.config/triage/agent.env`: ```bash TRIAGE_SERVER=https://triage.example.com TRIAGE_TOKEN=... ``` then start the agent: ```bash systemctl --user enable --now triage-agent.service ``` It runs `triage collect --follow --upload`. The first run reads the whole journal, then it follows new entries. The agent doesn't start until `agent.env` exists. ## Offline Events are stored on the host first, and only marked sent once the server accepts them. When the server can't be reached, the agent logs a warning and keeps collecting; it sends the backlog once the server is back. ## Hide more names Hosts redact their own hostname and every regular user's name. To hide other names, such as a GitHub account that shows up in repository URLs, add them to `agent.env`, comma-separated: ```bash TRIAGE_REDACT_NAMES=octocat,my-org ``` See [Privacy](/privacy) for everything that's redacted. ## Without a server `triage collect` on its own stores events on the host, and `triage issues` lists them. That's a good way to see what triage would send before you set anything else up. --- # Run a server Source: https://triage.timmo.dev/setup/server The server stores events from your hosts, groups them into issues and keeps the daily AI limits. Run one, somewhere every host can reach. It speaks plain HTTP on port 7171. For HTTPS, put it behind a reverse proxy such as Caddy, Traefik or nginx, or behind Cloudflare, and let that handle TLS. ## Choose where it runs The server itself is light, and runs on a Home Assistant Green. The models that decide on issues and suggest fixes need much more, and can run on another machine as a [worker](/setup/workers). See [Plan your hardware](/hardware) for what each part needs and the setups that work best. ## User service With the [Arch package](/install#arch-linux) installed, set any [options](/configuration#server) in `~/.config/triage/server.env`, then: ```bash systemctl --user enable --now triage-server.service loginctl enable-linger "$USER" # keep it running while you're logged out ``` The database lives in `~/.local/state/triage/server.db`, so `triage hosts add` and the other token commands work on it directly, without `--server`. ## Container ```bash docker compose up -d docker compose exec triage triage hosts add desktop ``` For HTTPS through Caddy, with certificates for your domain set up automatically: ```bash TRIAGE_DOMAIN=triage.example.com docker compose -f compose.yaml -f compose.caddy.yaml up -d ``` Or through a Cloudflare Tunnel, which needs no open ports at all. Create a tunnel in the Cloudflare dashboard, add a public hostname pointing at `http://triage:7171`, and use its token: ```bash TUNNEL_TOKEN=eyJ... docker compose -f compose.yaml -f compose.cloudflared.yaml up -d ``` The server always listens on `7171` inside the container, so proxies and tunnels can rely on it. Change the published port with `TRIAGE_PORT`. The database is in the `/data` volume. ## Home Assistant Install the [Home Assistant app](/install#home-assistant), then: 1. Set **server_admin_token** to a long random value, at least 32 characters, such as from `openssl rand -base64 32`, and start the app. 2. Add hosts and workers from any machine with `triage` installed, using `--server`, as below. The app's other options are **trust_proxy**, **decide_daily** and **suggest_daily**, which match the [settings](/configuration#server) of the same name. Small Home Assistant boards, such as the Green or Yellow, are rarely up to running models, so leave decide and suggest to a [worker](/setup/workers) on those. **Triage** in Home Assistant's sidebar opens the [issues page](/issues#in-a-browser) for Home Assistant's admins, with no admin token needed: Home Assistant has already signed them in. It's served on a separate port that only answers Home Assistant, so port 7171 still needs a token. The app stops briefly during backups so its database is copied consistently. Hosts keep their events until the server is back. ## Manage tokens remotely When you don't have a shell on the server, as with the Home Assistant app, manage tokens over the API. Set `TRIAGE_SERVER_ADMIN_TOKEN` on the server to a long random value, at least 32 characters, then from any machine: ```bash export TRIAGE_ADMIN_TOKEN= triage hosts add laptop --server https://triage.example.com triage workers add desktop --server https://triage.example.com ``` `add`, `list` and `remove` on `hosts`, `workers` and `admins` all take `--server`. `TRIAGE_SERVER_ADMIN_TOKEN` isn't stored or listed; you can leave it empty once you've added an admin of your own with `triage admins add`. ## Tokens | Token | Added with | Can | | --- | --- | --- | | Host | `triage hosts add ` | Send events | | Worker | `triage workers add ` | Fetch work and send back decisions and suggestions | | Admin | `triage admins add ` | Read issues, change their state and manage tokens | Each token is printed once, when it's added, and refused by the others' endpoints. The server only stores a hash of it. `list` shows who has a token, and `remove` revokes one. --- # Add workers Source: https://triage.timmo.dev/setup/workers A worker runs the models for a server. It fetches issues from the server, asks its models about them and sends back the answers, so the models only need to be reachable from the worker. The server could be a small always-on box while a desktop with a GPU does the work. ## Enrol On the server, or from anywhere with [`--server`](/setup/server#manage-tokens-remotely): ```bash triage workers add desktop ``` ## Run the worker On the worker, put the server's URL, the printed token and the model settings in `~/.config/triage/worker.env`: ```bash TRIAGE_SERVER=https://triage.example.com TRIAGE_WORKER_TOKEN=... TRIAGE_DECIDE=true TRIAGE_DECISION_MODEL=winnow ``` then start it: ```bash systemctl --user enable --now triage-worker.service ``` It runs `triage work`, which decides on new issues every 5 minutes with `TRIAGE_DECIDE`, and with `TRIAGE_SUGGEST` suggests fixes every 15 minutes for issues the configured decision model rated worth fixing with at least 0.8 probability. Turn on one or both. See [Configuration](/configuration#models) for the model settings and [Choices](/choices) for the models you can use. ## Limits The server's daily limits, `TRIAGE_DECIDE_DAILY` and `TRIAGE_SUGGEST_DAILY`, cover the server and every worker together, so adding workers doesn't raise them. When a worker is off, nothing queues up: the next one to run picks up the same issues. If the server or a model can't be reached, the worker logs a warning and tries again next time. ## On the server instead The server can run the same loops itself with `triage serve --decide` or `--suggest`, or `TRIAGE_DECIDE` and `TRIAGE_SUGGEST` in `server.env`, when the models are reachable from there.