How 3AM decides
Models
What the optional on-prem model does, where it runs, and what it is never used for.
3AM can run a small language model inside your network to write the incident note for the people on call. That is all it is used for.
| Profile | Model | Runs on |
|---|---|---|
| CPU | qwen2.5:1.5b-instruct | Any server with 4 GB+ free memory; a note in a few seconds |
| GPU | qwen2.5:7b-instruct | NVIDIA (8 GB+ VRAM) or AMD ROCm |
| None | Incident notes use a template |
The installer picks the profile from the hardware (--profile overrides it). The model runs in its own container,
reachable only by 3AM, with its weights loaded from the bundle.
Guardrails
- The model sees only the incident's verified facts: the confirmed cause, the evidence, the proposed change, the policy decision. It sees nothing it could invent from.
- A note that leaves out the confirmed cause or the proposed command is discarded, and the template is used.
- Every call is recorded: the model, a hash of the prompt and a redacted copy, the output, tokens and time taken.
- Diagnosis, the choice of fix and every approval happen without the model. Removing it changes only how the note reads.
Your own model endpoint
To use a model server you already run (any OpenAI-compatible endpoint, such as vLLM or Ollama), set on the 3AM container:
| Variable | Example |
|---|---|
THREEAM_MODEL_URL | http://models.bank.internal:8000/v1 |
THREEAM_MODEL | qwen2.5-7b-instruct |
THREEAM_MODEL_KEY | A bearer token, if your endpoint needs one |