Self-hosting
One binary, one data directory. Everything you need to run, secure and back up open-server, the free VAKYN server, on your own machine.
open-server is open source (MIT) on GitHub. It serves the same API as VAKYN MAX, so code written against one runs against the other once you change the base URL and the key. It brings its own dashboard for keys, usage and a playground, and a server page that shows its status.
Install
| Method | Where |
|---|---|
| One-line installer (curl on macOS and Linux, PowerShell on Windows) | Below; it picks the right build and checks it |
| Prebuilt binaries (macOS Apple Silicon, Linux x64 and arm64, Windows x64) | On every GitHub release; see Release archives |
| Docker images (CPU and CUDA) | On GHCR with every release; see Docker |
| Build from source | See Building from source |
# Linux and macOS (Apple Silicon)curl -fsSL https://github.com/vakyn-ai/open-server/releases/latest/download/install.sh | sh # Windows (PowerShell)irm https://github.com/vakyn-ai/open-server/releases/latest/download/install.ps1 | iexinstall.sh picks the build for your system: on Linux x64 with an NVIDIA GPU and driver R570 or newer, the CUDA build for the GPU's family (from nvidia-smi); otherwise the CPU build, and it says why. It downloads it from the latest release, checks its SHA-256 and installs vakyn to ~/.local/bin (/usr/local/bin as root). On Linux with systemd it then asks Install vakyn as a service and start it? [Y/n]: yes by default on a terminal, no when there is none. install.ps1 installs vakyn.exe to %LOCALAPPDATA%\vakyn\bin and adds it to your user PATH. Running an installer again upgrades in place and restarts an installed service.
| install.sh flag | Environment | Meaning |
|---|---|---|
--version 0.2.0 | VAKYN_VERSION | Install this release instead of the latest (install.ps1 too) |
--dir DIR | VAKYN_INSTALL_DIR | Where the vakyn command goes (install.ps1 too) |
--cpu | VAKYN_CPU=1 | The CPU build even with an NVIDIA GPU |
--target T | VAKYN_TARGET | A given archive target, e.g. linux-x64-cuda-sm90 |
--service | VAKYN_SERVICE=yes | Install and start the service without asking |
--no-service | VAKYN_SERVICE=no | Don't install the service |
VAKYN_SERVICE_ARGS | vakyn serve flags for the service, e.g. --model /srv/VAKYN-4B-UQFF | |
GITHUB_TOKEN / GH_TOKEN | Download with this token (a private fork, or GitHub API rate limits) |
Flags go after sh -s --, for example curl -fsSL …/install.sh | sh -s -- --version 0.2.0 --no-service. The Linux builds need glibc 2.35 or newer; on Alpine and other musl systems use the Docker image.
Release archives
Each release has one archive per target, named vakyn-<version>-<target>, with a .sha256 file next to it and all checksums in SHA256SUMS. Every binary includes the model engine.
| Target | For | Engine |
|---|---|---|
linux-x64-cpu | Linux x86-64 with glibc 2.35+ (Ubuntu 22.04, Debian 12 and newer) | CPU |
linux-arm64-cpu | Linux arm64 with glibc 2.35+ | CPU |
linux-x64-cuda-sm80 | NVIDIA Ampere and Ada: A100, A10, A40, RTX 30xx, RTX 40xx, L4, L40S | CUDA, flash attention |
linux-x64-cuda-sm90 | NVIDIA Hopper: H100, H200 | CUDA, flash attention |
linux-x64-cuda-sm100 | NVIDIA Blackwell data center: B200, GB200 | CUDA, flash attention |
linux-x64-cuda-sm120 | NVIDIA Blackwell RTX: RTX 50xx, RTX PRO 6000 | CUDA, flash attention |
macos-arm64-metal | macOS on Apple Silicon | Metal |
windows-x64-cpu | Windows x64 (.zip) | CPU |
Flash attention is compiled for one GPU architecture at a time, so there is a CUDA build per GPU family. nvidia-smi --query-gpu=compute_cap --format=csv shows yours: 8.x is sm80, 9.0 is sm90, 10.0 is sm100, 12.0 is sm120. The CUDA builds use CUDA 12.8, need an NVIDIA driver R570 or newer, and carry the CUDA runtime libraries they use in lib/ next to the binary, so no CUDA toolkit is needed.
tar -xzf vakyn-<version>-linux-x64-cpu.tar.gzsha256sum -c vakyn-<version>-linux-x64-cpu.tar.gz.sha256./vakyn-<version>-linux-x64-cpu/vakyn serveThe web app (front page, server page and dashboard) is built into the binary; nothing else needs to be installed or served at runtime.
Building from source
You need Rust (stable) and Node.js 22 for the web app that is embedded in the binary:
git clone https://github.com/vakyn-ai/open-server.gitcd open-server(cd web && npm ci && npm run build)cargo build --release# the binary: ./target/release/vakyn — put it on your PATH./target/release/vakyn --versionCommands
| Command | What it does |
|---|---|
vakyn serve | Run the server in the foreground (logs to stderr; Ctrl-C stops it) |
vakyn pull REPO [--quant Q6K] | Download a release from Hugging Face into the cache without serving it; see Downloading releases |
vakyn start | Run the same server in the background, with the same flags; logs go to <data-dir>/vakyn.log |
vakyn status | Show whether it is running, with pid, address and uptime (exit code 3 when it is not) |
vakyn stop [--timeout 30] | Stop the background server gracefully; kill it if it has not stopped after the timeout |
vakyn keys create --name NAME | Create an API key and print it once |
vakyn keys list | List keys: name, prefix, seed, created, last used, revoked |
vakyn keys revoke NAME_OR_PREFIX | Revoke keys by name or by prefix (vk_ plus 8 characters) |
vakyn db backup [--out PATH] | Write a consistent copy of the database; safe while the server runs |
vakyn db restore PATH | Replace the database with a backup (server stopped; the current database is backed up first) |
vakyn db reset [--yes] | Delete the admin account, keys and usage after taking a backup (server stopped) |
vakyn service install [FLAGS] | Install vakyn as a systemd service, enable it at boot and start it (Linux); see Service |
vakyn service uninstall [--purge] | Stop and remove the service; --purge also deletes its data directory |
vakyn --version | Print the version |
Every command accepts --data-dir, so several servers can run side by side on different ports and directories. Keys can also be created, listed and revoked on the dashboard's API keys page after signing in.
Per-key seed
Each API key gets a random 64-bit inference seed when it is created (keys from earlier versions get one when you upgrade). It never changes, even after the key is revoked, and every request made with the key passes it to the engine. Playground runs use one fixed seed of their own. The seed is shown read-only on the API keys page and by vakyn keys list.
Today the model computes its answers without sampling, so the seed does not change any numbers: the same request gets the same probabilities with any key. The seed applies wherever the engine uses randomness.
Service (Linux)
vakyn service install --model /srv/VAKYN-4B-UQFF # user servicesudo vakyn service install --model /srv/VAKYN-4B-UQFF # system servicevakyn service install --dry-run --port 9000 # print the unit onlyvakyn service uninstall # keeps the datavakyn service uninstall --purge # deletes it too- With
sudo(or as root) it writes/etc/systemd/system/vakyn.service, which runsvakyn serveas a dedicatedvakynuser on/var/lib/vakyn. Otherwise it writes a user unit in~/.config/systemd/user/on your own data directory; a user service runs while you are logged in, andsudo loginctl enable-linger $USERstarts it at boot.--userand--systemchoose explicitly. - Flags after
installarevakyn serveflags for the service; they are checked, and relative paths are made absolute.--model vakyn/VAKYN-4B-UQFFworks too: the service downloads the release on its first start (putHF_TOKEN=...in the env file below for a private release;--hf-tokenis refused so the token never lands in the unit). An--admin-password(orVAKYN_ADMIN_PASSWORD) is applied at install time and never written into the unit; a generated admin password is printed once byservice install. Other settings (RUST_LOG, ...) can go in/etc/vakyn/vakyn.envor~/.config/vakyn/vakyn.env. - Installing again with other flags updates the unit and restarts the service; with the same flags nothing changes. It refuses while
vakyn startruns on the same data directory. - Status and logs:
systemctl status vakynandjournalctl -u vakyn -f(add--userfor a user service). API keys for the system service:sudo -u vakyn vakyn keys create --name app --data-dir /var/lib/vakyn. - macOS and Windows services are not supported. There,
vakyn servicesays so, andvakyn startruns VAKYN in the background.
Server flags
For vakyn serve and vakyn start:
| Flag | Environment | Default | Meaning |
|---|---|---|---|
--host | VAKYN_HOST | 127.0.0.1 | Address to listen on. Use 0.0.0.0 to accept connections from other machines |
--port | VAKYN_PORT | 8080 | Port; 0 picks a free one (printed at start) |
--data-dir | VAKYN_DATA_DIR | per OS, below | Where the database, log and backups live |
--admin-user | VAKYN_ADMIN_USER | admin | Dashboard username |
--admin-password | VAKYN_ADMIN_PASSWORD | generated once | Dashboard password; when given, it is set on every start |
--max-queue | 32 | Requests admitted at once, running or waiting; more get 429 | |
--session-ttl | 604800 (7 days) | Dashboard sign-in lifetime, in seconds | |
--login-attempts | 5 | Failed sign-ins from one address before it has to wait (1 s, then doubling, up to 15 minutes); 429 with retry-after meanwhile | |
--login-global-attempts | 50 | Failed sign-ins across all addresses before everyone has to wait (up to 1 minute) | |
--trust-proxy | off | Take the client address for sign-in limiting from X-Forwarded-For; turn on only behind a reverse proxy | |
--site-url | VAKYN_SITE_URL | https://vakyn.com | The VAKYN website the web app links to for docs, skills and the FAQ |
--engine | mistralrs with --model, else fake | Which engine answers: the model (see Model and tuning) or a deterministic test engine | |
--fake-latency-ms | 0 | Test engine only: delay added to every request | |
--log | Also write the server log to this file |
Log verbosity follows RUST_LOG (default info). At start the server logs its effective settings; the public server page (/server on your server) shows the same values.
Model and tuning
VAKYN releases are on Hugging Face, for example vakyn/VAKYN-4B-UQFF. Name one with --model and the server downloads it on first start, or point --model at a release folder you already have:
vakyn serve --model vakyn/VAKYN-4B-UQFF # downloads, then serves Q8_0vakyn serve --model VAKYN-4B --quant Q4K # short name; or Q6Kvakyn serve --model ./VAKYN-4B-UQFF # a local release folder# or with a config file holding the same keys (flags win)vakyn serve --config vakyn.toml| File | What it is |
|---|---|
VAKYN-4B-Q8_0-0.uqff | The weights at 8 bits: the default, closest to the reference answers |
VAKYN-4B-Q6K-0.uqff | 6 bits: smaller, answers stay close; --quant Q6K |
VAKYN-4B-Q4K-0.uqff | 4 bits: smallest, answers drift more; --quant Q4K |
config.json, residual.safetensors | What the engine needs next to the UQFF files |
vakyn_head.safetensors | The pointer head that scores the options |
calibration.json | Temperatures that calibrate the probabilities |
tokenizer.json | The tokenizer |
The server downloads and loads only the quantization you choose.
The model loads in the background: the server page (/server) shows loading and the API answers 503 with retry-after until it is ready. If the model cannot load (for example the context does not fit in GPU memory), the server stops with a message saying what to change. The effective settings are logged at start and shown on the server page.
| Flag / config key | Default | Meaning | llama-server analogue |
|---|---|---|---|
--ctx-size | 65536 | KV cache size in tokens (the default fits the largest request the API accepts) | -c |
--gpu-memory-fraction | Size the KV cache as a share of GPU memory instead | ||
--gpu-memory-mb | Size the KV cache in MB instead | ||
--cache-type | auto | f8e4m3 halves the KV cache (CUDA) | --cache-type-k/v |
--max-seqs | 32 | Sequences the engine runs at once | -np |
--max-batch-tokens | engine default | Tokens per scheduler step | -b |
--prefill-chunk | engine default | Prompt tokens per prefill chunk | -ub |
--prefix-cache-n | 6 | Documents kept for reuse; each cached 32k-token document holds about 1.2 GB of GPU memory (0 disables) | --cache-reuse |
--paged-attn | auto (off) | Paged attention; off by default because it is slower here, see the note below | |
--cpu | off | Run on the CPU even with a GPU | |
--device-layers | automatic | Layers per GPU: 0:20 1:12, or one number for GPU 0; layers not placed run on the CPU (same answers, several times slower) | -ngl |
--dtype | auto | Activations: f16 on GPUs, the engine's choice on CPU | |
--load-timeout | 900 | Seconds for loading, warm-up and the fit check before startup fails with the step it was on (0 waits forever) | |
--seed | Seed for the engine's startup checks; API and playground requests use their key's seed (see Per-key seed) | -s |
Each request reads its document once and then answers all its questions in one batched pass; requests from different clients run side by side. On a GPU the server checks at startup that --prefix-cache-n + 1 maximum-length documents fit, so an oversized setting fails then and not under load. Before loading on a single NVIDIA GPU, it also estimates the memory the model needs at --ctx-size. If the GPU has less free, startup fails with that number and what fits (a lower --ctx-size or --prefix-cache-n, --quant Q4K or a smaller model) instead of quietly running part of the model on the CPU; --device-layers N runs the remaining layers on the CPU on purpose. Requests longer than --ctx-size (or the API's 32,768-token limit) get the same 400 as any over-limit request.
Downloading releases
export HF_TOKEN=hf_... # private releases onlyvakyn pull vakyn/VAKYN-4B-UQFF # download Q8_0, print the foldervakyn pull VAKYN-4B --quant Q6K # another quantvakyn serve --model vakyn/VAKYN-4B-UQFF # serves from the cache, no network- What is downloaded: the weights of one quant and the files the server needs next to them (
config.json,residual.safetensors,tokenizer.json,vakyn_head.safetensors,calibration.json,generation_config.json,chat_template.jinja,manifest.json). Other quants are not. - Token:
--hf-token, else theHF_TOKENenvironment variable, else the token saved byhf auth login. It is never printed or logged. Withvakyn start, preferHF_TOKEN: flags show up on the process command line. - Cache: the standard Hugging Face cache (
~/.cache/huggingface/hub, orHF_HUB_CACHE/HF_HOME/hub), shared withhf download.--model-dir DIR(VAKYN_MODEL_DIR, ormodel_dirin the config file) keeps the files inDIRinstead, with the same layout. - Progress: a bar per file and one for the total on a terminal; a log line every 10 seconds under Docker or systemd. While it downloads, the server already listens:
GET /api/statusreports"status": "downloading"withdownload.bytes_doneanddownload.bytes_total, the server page shows "Downloading 1.2 / 4.5 GB", and the API answers503withretry-after. - Resume and checks: an interrupted download resumes where it stopped (run the same command again), and every file is checked against the SHA-256 in the release's
manifest.json; a file that doesn't match is downloaded again. - Offline: once a quant is complete in the cache,
vakyn serveuses it without the network.vakyn pullalways checks Hugging Face for a newer release. - Errors say what to do: a private release without a token (set
HF_TOKENor runhf auth login), an unknown release or quant (the available quants are listed), a full disk, or a lost connection (retried with back-off; then run the command again to resume).
Building with the model engine
Prebuilt releases include the engine. From source, choose the backend with a cargo feature:
(cd web && npm ci && npm run build)cargo build --release --features cuda,flash-attn # NVIDIA (set CUDA_COMPUTE_CAP, e.g. 120 for RTX 50xx)cargo build --release --features metal # Apple Siliconcargo build --release --features engine-mistralrs # CPU onlyFlash attention matters on CUDA: long documents are about three times slower without it.
Data directory
| System | Default |
|---|---|
| Linux | ~/.local/share/vakyn |
| macOS | ~/Library/Application Support/ai.vakyn.vakyn |
| Windows | %APPDATA%\vakyn\vakyn\data |
| File | Contents |
|---|---|
vakyn.db | SQLite: the admin account (password hash), API keys (hashes and seeds), usage events, dashboard sessions, the playground seed |
vakyn.lock | Held by the running server so two servers never share a directory |
vakyn.pid | The running server's pid and address, for status and stop |
vakyn.log | Output of vakyn start |
backups/ | Copies written by db backup, db restore and db reset |
Keys and passwords are stored only as argon2 hashes; a key is shown once, when it is created. Documents and answers are not stored, only per-request counts (time, key, input tokens, status, request id).
Backup and restore
# while the server runsvakyn db backup # -> <data-dir>/backups/vakyn-YYYYMMDD-HHMMSS.dbvakyn db backup --out /mnt/backup/vakyn.db # restore: stop firstvakyn stopvakyn db restore /mnt/backup/vakyn.db # the current database is backed up firstvakyn startvakyn db reset asks for confirmation (skip it with --yes), backs up, then deletes everything; the next start creates a new admin account. Back up before upgrading.
Network and HTTPS
- By default the server only listens on
127.0.0.1, so only the same machine can reach it. - There is no built-in TLS. To serve other machines, put a reverse proxy that terminates HTTPS in front (Caddy, nginx, Traefik) and keep VAKYN on localhost behind it.
- Behind a proxy every request comes from the proxy's address, so start VAKYN with
--trust-proxy: failed sign-ins are then counted per client, using the lastX-Forwarded-Forentry (the one the proxy adds). Without a proxy leave it off, since clients could forge the header. - The front page and the server page are public; the dashboard needs the admin sign-in; the API needs a key.
Example: Caddy
vakyn.example.com { reverse_proxy 127.0.0.1:8080}Docker
Every release pushes one image per engine build to the GitHub Container Registry. All of them run as a non-root user (uid 10001), listen on 0.0.0.0:8080 and keep the database and the model cache in the /data volume (/data/.cache/huggingface/hub), so a restart doesn't download the model again.
| Tag | For |
|---|---|
cpu | Any x86-64 or arm64 machine, no GPU |
nvidia-sm60 | Pascal: P100 |
nvidia-sm70 | Volta: V100 |
nvidia-sm75 | Turing: T4, RTX 20xx |
nvidia-sm80 | Ampere and Ada: A100, A10, RTX 30xx, RTX 40xx, L4, L40S |
nvidia-sm90 | Hopper: H100, H200 |
nvidia-sm100 | Blackwell data center: B200, GB200 |
nvidia-sm120 | Blackwell RTX: RTX 50xx, RTX PRO 6000 |
Each tag also exists per release, for example 0.1.0-nvidia-sm70. To find your GPU family, run nvidia-smi --query-gpu=compute_cap --format=csv: 8.6 means nvidia-sm80, 7.0 means nvidia-sm70. A CUDA image refuses to start on another family and names the right tag. Images from sm80 up use flash attention.
docker run -d --name vakyn --gpus all -p 8080:8080 -v vakyn-data:/data \ -e VAKYN_MODEL=vakyn/VAKYN-4B-UQFF -e VAKYN_QUANT=Q8_0 \ ghcr.io/vakyn-ai/open-server:nvidia-sm80 # A release folder you already have, mounted under /modelsdocker run -d --name vakyn -p 8080:8080 -v vakyn-data:/data \ -v "$PWD/VAKYN-4B-UQFF:/models/VAKYN-4B-UQFF:ro" -e VAKYN_MODEL=/models/VAKYN-4B-UQFF \ ghcr.io/vakyn-ai/open-server:cpu docker logs vakyn # the admin password generated on first startdocker exec vakyn vakyn keys create --name app # create an API key- Every
vakyn serveflag has aVAKYN_*environment variable:--ctx-sizeisVAKYN_CTX_SIZE,--max-seqsisVAKYN_MAX_SEQS, and so on. The full list is in BUILDING.md in the repository. - Set the admin password with
-e VAKYN_ADMIN_PASSWORD=…, or read the generated one fromdocker logson the first start. Pass-e HF_TOKENfor private models. - A named volume (
vakyn-dataabove) needs no setup. A host directory mounted at/datamust be writable by uid 10001 (sudo chown 10001 /srv/vakyn). - The CUDA images need the NVIDIA Container Toolkit and driver R570 or newer (R580 for Pascal).
- The port is published on all host addresses by
-p 8080:8080; use-p 127.0.0.1:8080:8080to keep it local behind a reverse proxy (with--trust-proxy).
Building the images
docker build -t vakyn:cpu .docker build --build-arg VARIANT=cuda --build-arg CUDA_COMPUTE_CAP=70 -t vakyn:nvidia-sm70 .A CUDA build compiles the kernels for the one GPU family you choose: tens of minutes, or hours from sm80 up because of flash attention. BUILDING.md in the repository covers every platform.