Recorded by CI from a real session, on every push.
Running a local LLM still means picking an engine, matching a model to your VRAM, choosing a quantization and wiring up a service. SibillaOS does all of that for you. Boot the ISO, walk through the standard Ubuntu screens (locale, network, disk, your user), and the first thing your machine does after installing is serve an OpenAI-compatible API.
$ curl http://myserver:8080/v1/chat/completions \
-H "Authorization: Bearer $(sudo cat /etc/llmd/apikey)" \
-d "{\"model\": \"$(sudo cat /etc/llmd/model)\", \"messages\": [{\"role\": \"user\", \"content\": \"hello\"}]}"Already on Ubuntu 24.04? You do not have to reinstall anything. Add the repository and turn the box into an LLM appliance:
$ curl -fsSL https://engineering87.github.io/sibillaos/apt/sibillaos-archive-key.asc \
| sudo gpg --dearmor -o /usr/share/keyrings/sibillaos-archive-keyring.gpg
$ printf 'Types: deb\nURIs: https://engineering87.github.io/sibillaos/apt/\nSuites: ./\nSigned-By: /usr/share/keyrings/sibillaos-archive-keyring.gpg\n' \
| sudo tee /etc/apt/sources.list.d/sibillaos.sources
$ sudo apt update && sudo apt install llmd
$ sudo sibilla setupsibilla setup detects the hardware, installs the engine if it is missing, pulls a fitting model and serves the API on port 8080. It is a good guest: it does not touch your firewall or your other package sources, and trying it is reversible: sudo sibilla remove takes out everything it installed, restores what it displaced and leaves the machine as it found it (CI verifies exactly that on every push).
For a whole machine or a VM from scratch, install one of the images instead. Two paths, both end at the same place.
Cloud image (fastest, no USB). Download the qcow2 for your architecture from Releases, verify it against SHA256SUMS-cloud-<arch>, and boot it in your hypervisor with a cloud-init user-data that sets your user and SSH key, exactly as you would any Ubuntu cloud image.
ISO (bare metal or VM). Download the .part files and SHA256SUMS from Releases, reassemble and verify, write to a USB drive, boot it and pick "Install SibillaOS (automated)":
$ cat sibillaos-*-amd64.iso.part* > sibillaos.iso
$ sha256sum -c SHA256SUMS
$ sudo dd if=sibillaos.iso of=/dev/sdX bs=4M status=progressEither way, first boot detects the hardware, picks the engine and a fitting model, downloads it and brings up the API. Then, on the machine:
$ sibilla status
engine: ollama (active)
model: hf.co/bartowski/Qwen_Qwen3-4B-GGUF
api: http://192.168.1.50:8080 (key in /etc/llmd/apikey)
$ KEY=$(sudo cat /etc/llmd/apikey)
$ MODEL=$(sudo cat /etc/llmd/model)
$ curl http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$MODEL\", \"messages\": [{\"role\": \"user\", \"content\": \"hello\"}]}"From another machine, swap localhost for the address sibilla status prints and copy the key from /etc/llmd/apikey.
Where to go next:
- switch model:
sudo sibilla model use ID(sibilla model listshows what fits your machine) - HTTPS:
sudo sibilla tls enable myserver.lan(add--acme you@example.orgfor a public hostname) - editor and agent config:
sudo sibilla connect(--writeplaces Continue and aider configs;--env,--mcp,--snippet python|nodeemit ready-made formats;--remotegenerates the script that configures your workstation, key prompted there) - agents via MCP:
sudo sibilla mcp enableserves the model as Model Context Protocol tools - chat interface:
sudo sibilla webui enable(Open WebUI on port 3000) - something wrong:
sudo sibilla doctorproduces a paste-ready report for your issue, secrets excluded - how fast is my machine:
sudo sibilla benchmeasures first-token latency and tokens/s, in a table worth sharing
The installer detects your hardware and makes the decisions a human would otherwise have to research.
| Engine selection | vLLM in an OCI container on datacenter GPUs (24 GB VRAM and up), Ollama everywhere else, CPU-only machines included. |
| Model sizing | llmfit recommends only models that actually fit your VRAM and RAM, with the best quantization and a speed estimate. |
| Curated catalog | Permissively licensed (Apache-2.0/MIT), non-gated Hugging Face repos only. Verified ids, signed list. |
| Resilient download | The model is pulled from Hugging Face during install and resumed at first boot if the connection drops. |
| Single endpoint | One OpenAI-compatible API on port 8080 with mandatory bearer tokens: multiple keys with per-key revocation (sibilla key), structured access logs. Engines stay on loopback. sibilla tls enable HOSTNAME switches the gateway to HTTPS (local CA, or Let's Encrypt with --acme). |
| One CLI | sibilla status is a health view of the whole stack: engine, served models, disk usage of the model store, GPU utilization, gateway reachability. |
| Model management | sibilla model list shows what fits your machine, sibilla model use ID downloads and switches the served model, import FILE brings one in from a USB stick or shared drive (accepted only if it matches the signed catalog digests), rm and prune reclaim disk. |
| Observability | sibilla metrics enable serves Prometheus metrics behind the same API key; Grafana dashboard included in docs/observability. |
| Chat interface | sibilla webui enable starts Open WebUI on port 3000 as an opt-in container, wired to the local engine. |
| Editor hookup | sibilla connect prints ready-to-paste configuration for VS Code (Continue, Cline), aider and any OpenAI-compatible client. |
| Agent hookup | sibilla mcp enable exposes the local model to MCP clients (Claude Code and other agent frameworks) as chat and list_models tools, behind the same API keys. See docs/mcp.md. |
| Config as code | Declare the desired state (model, TLS, metrics, MCP, WebUI) in one profile file; sibilla apply converges the machine onto it, idempotently. apply export turns a configured machine into a profile; cloud-init or a fleet tool drops the file and first boot picks it up. See docs/configuration.md. |
| Air-gapped | Machines with no outbound network install from a companion payload volume: models travel by USB stick, verified against the signed catalog digests on both ends. CI proves it on every push in a network-restricted VM. See docs/airgap.md. |
| Local RAG | /v1/embeddings through the same gateway and keys: pull a catalog embedding model (sibilla model pull) and point any OpenAI-compatible RAG framework at this machine. See docs/embeddings.md. |
The base install is a headless server; a desktop variant is on the roadmap.
Fair question: curl ollama.com/install.sh | sh is one line, and SibillaOS ships that very engine. The difference is everything around it, and it matters the moment the machine serves anyone but you:
| Plain engine install | SibillaOS |
|---|---|
| API open to whoever reaches the port | Mandatory bearer keys from the first second, per-client keys with instant revocation, TLS in one command |
| You trust whatever the download gave you | Models verified against a GPG-signed catalog with per-file sha256 digests; a mismatched artifact is refused |
| Configuration lives in your shell history | One declared profile file; sibilla apply converges any machine onto it, cloud-init included |
| Needs the internet to exist | Installs and serves fully air-gapped from a verified USB payload |
| Uninstall is a guess | sibilla remove takes out exactly what was installed and restores what it displaced, CI-verified |
| You wire up agents, RAG and editors by hand | MCP tools, /v1/embeddings and editor configs served by the same box, behind the same keys |
None of this is a criticism of Ollama, which does its job excellently - it is the job of turning an engine into an appliance someone can audit, replicate and walk away from. If you serve only yourself on localhost, plain Ollama is probably all you need. The moment there is a network, a team, a compliance question or a second machine, the difference is the product.
The engines never listen off-host: the gateway is the only door in. The full decision log is in docs/architecture.md.
Prebuilt images are the quick path above; to build your own you only need xorriso and curl (no root). The build downloads the official Ubuntu 24.04 live-server ISO, verifies its checksum and repacks it with the SibillaOS autoinstall, packages and branding:
sudo apt-get install xorriso
./packages/build-debs.sh # build the llmd-* debs
./iso/build.sh # out/sibillaos-<version>-amd64.isoEvery push to main also builds the ISO in CI, boots it in QEMU and runs the automated install end to end; the image is attached to each run as an artifact.
Installed systems receive llmd package updates through the project's signed APT repository, preconfigured on every image: a plain sudo apt update && sudo apt upgrade keeps the stack current between reinstalls.
iso/ ISO repack (official Ubuntu live-server + payload) and autoinstall
cloud/ qcow2 cloud image bake (official Ubuntu cloud image + payload)
packages/ Debian packages: hardware detection, engines, gateway, first boot
catalog/ curated model list (signed JSON)
branding/ logo, banner and wallpaper
docs/ architecture document
Working proof of concept: on every push, CI builds the ISO, boots it under BIOS and UEFI, runs the install end to end and gets a real chat completion through the gateway on first boot. Engine versions are pinned in the installer. Release ISOs walk you through the standard installer screens and you choose your own credentials; only the fully unattended CI images use a fixed test user. The design and decision log live in docs/architecture.md; where the project is going is in ROADMAP.md.
CI proves everything except silicon: if you own an NVIDIA card, an AMD
card or a Ryzen APU, thirty minutes and docs/validation/gpu.md
turn your machine into exactly the data this project needs. sibilla bench and sibilla doctor produce paste-ready, secret-free results.
Contributions are welcome. CONTRIBUTING.md covers the build setup, the CI test suite your change has to pass, and the criteria for model catalog additions.
Apache-2.0, see LICENSE. Bundled components keep their own licenses: vLLM (Apache-2.0), Ollama (MIT), llmfit (MIT). NVIDIA drivers are not redistributed by this repository; the ISO installs them from the Ubuntu restricted component.
