Ollaya – Ollama for open-source, Jev-style decision models
Ask typed questions about any text or JSON and get calibrated answers in milliseconds. Private, open source, on your own hardware.
Decisions in milliseconds.
A decision model answers in a single forward pass, with no token-by-token generation. On your own GPU, a five-question request to Laya takes about 10 ms, end to end through the HTTP API.
RTX 4090, five questions, end to end
Hosted API, median request
Ollaya: median of a five-question request through the HTTP API on an NVIDIA RTX 4090 (laya in fp16, the others in fp32). Jev: median request latency of the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), which includes the network. Setups differ, so read it as an order-of-magnitude comparison.
Speaks TypeSafe’s API.
Ollaya serves /v1/systemone and /v1/models with TypeSafe’s request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.
TypeSafe compatibility guide
Open weights, ready to pull.
Pick by what you need: laya is the fastest, decider the most accurate, von reads up to 8,192 tokens, and qwen3guard screens text for safety. The models page shows each one’s accuracy and speed.
Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.
Runs on your machine with ONNX Runtime, on the CPU or an NVIDIA GPU. The server listens on 127.0.0.1 by default.
Weights come from their authors’ Hugging Face repositories, pinned to a commit and checked against sha256. Ollaya never re-hosts them, and the runtime is Apache-2.0.
Run as many decisions as your hardware can handle. No metering and no API bill.
Probabilities you can put thresholds on. Each model ships its own calibration, and a Modelfile refits it on your labelled data.
A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, Windows, WSL 2 or Docker takes a request down to milliseconds.
NVIDIA GPUs need driver R580 or newer; the install scripts fetch the CUDA libraries only when they find one. On a Mac, laya and nli run on the Apple GPU through MLX; other models, AMD and Intel GPUs, and the Windows and Linux desktop apps use the CPU.
One binary, one command: ollaya run laya.
macOS, Windows, Linux and Docker · Apache-2.0 · GitHub