2026-10-11 17:11 UTC

Janus's maintainer claims the released single-Go-binary server runs GGUF models via llama.cpp's Vulkan backend across AMD/Intel/NVIDIA with an OpenAI-compatible API and no Python/Docker/Ollama dependencies; sustained external adoption as a practical cross-vendor CUDA-free local inference option confirms it, stagnation marks another modest Show HN release.

state: resolvedheat: lowuncertainty: lowknownscott: lowlocal-inference inference-runtimes vulkanVibra-Ingenn

What is this?

Per the submission itself, Janus is a Show HN release by Vibra-Ingenn: a single Go binary that serves GGUF models through llama.cpp's Vulkan backend on AMD/Intel/NVIDIA GPUs, exposing an OpenAI-compatible API with no Python, Docker, or Ollama dependency. Notably, the supplied web results contain no direct mention of Janus or its maintainer, so the project's claims, quality, and adoption cannot be verified from the snippets. What the results do establish is the competitive context: llama.cpp already ships a first-class Vulkan backend and an OpenAI-compatible llama-server, and the thin-single-binary-wrapper niche is crowded โ€” Shimmy (~5.9k stars) pitches a 5MB Rust OpenAI-compatible GGUF server, and ollama-rdna1 exists specifically to serve ROCm-abandoned AMD cards via llama.cpp+Vulkan. Whether Janus adds anything over plain llama-server, and whether it sustains adoption, is the open question.

Why it matters to Scott

Scott's own wikis already carry this operating pattern: dev:technology.ollama runs local GGUF serving behind an OpenAI-compatible endpoint (gamepc:11434) with LiteLLM as the gateway, and Janus repeats that canon shape (lean server, OpenAI API, no Docker) rather than extending it โ€” its one new axis, cross-vendor Vulkan without CUDA, targets hardware his WSL2/CUDA gamepc stack doesn't run. It only upgrades from LOW if adoption (the case's own resolvable question) proves it a meaningfully leaner or more trustworthy Ollama swap โ€” e.g. if the pending Ollama silent-context-truncation case turns bad โ€” or if it differentiates itself from plain llama-server, which already ships a Vulkan backend and an OpenAI-compatible API.
dev:technology.ollamadev:project.gamepcdev:technology.litellmdev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:concept.llama-cppradar:concept.ggufradar:concept.ollamaradar:lemonade-vulkan-rocm-dropradar:panther-lake-vulkan-backend-comparisonradar:llamacpp-fork-fragmentation
queries asked of Scott's wikis
  • local inference Vulkan CUDA-free AMD GPU runtime
  • OpenAI-compatible API shim over llama.cpp
  • single-binary LLM server deployment friction no Docker
  • local models inside coding agent harness
  • Ollama alternative dependency-free GGUF serving
  • GGUF local models for RAG and knowledge systems

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

09-30 14:00โญ origin echo-reconstructedREADME: "Janus is a single Go binary that runs .gguf models on your machine (GPU or CPU) and exposes an OpenAI-compatible API. No Python, no
Vibra-Ingenn (GitHub org; HN submitter Maverick617, commits authored as "Janus Cleanup" <[email protected]>, co-authored with Claude Opus 4.6) on github (echo) ยท attributed from hn.story.49926773
โ€”
10-01 20:36first on hacker news ยท published ยท +30.6hShow HN: Janus โ€“ Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
Maverick617
โ€”
10-01 20:36amplified on hacker news ๐Ÿ‘‘hn.story.49926773
Maverick617
peak 84 ยท 14 comments ยท 100% of case engagement
10-01 23:21our radar first saw it ยท +33.4hdiscovery anchor: hn.story.49926773โ€”

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: Janus โ€“ Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
Retrieved article excerpt

Open article ยท Retrieved 2026-10-01T23:31:33.483298+00:00

# Janus โ€” Local LLM Server & OpenAI-Compatible API

Janus is a **single Go binary** that runs `.gguf` models on your machine (GPU or CPU) and exposes an **OpenAI-compatible API**. No Python, no Docker, no Ollama required.

**Use it your way:** call it from the **command line** (`curl`, PowerShell, scripts), wire it into **Cursor / Cline / any OpenAI client** โ€” same local models, whatever workflow fits you.

---

## What you get

- **Local inference** โ€” llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
- **OpenAI-compatible API** โ€” `/v1/chat/completions`, `/v1/models`
- **Hot-swap models** โ€” change `.gguf` without restarting
- **Thinking model support** โ€” `<think>` reasoning split into `reasoning_content`
- **Chat template auto-detection** โ€” uses the template from GGUF metadata
- **Zero dependencies** โ€” one `.exe` on Windows, no Python, no Docker

---

## Requirements

| Platform | What you need |
| --- | --- |
| **Windows** (primary) | Windows 10/11, [Go 1.22+](https://go.dev/dl/), Vulkan-capable GPU recommended |
| **Linux** | Go 1.22+, Vulkan or CPU |
| **macOS** | Go 1.22+, CPU backend (Vulkan varies by hardware) |

**Disk:** plan for the model size (often 2โ€“8 GB per model) plus ~50 MB for Janus + llama.dll.

---

## Quick start (Windows)

### 1. Clone and build

```
git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
.\build.ps1
```

`build.ps1` downloads pre-built **llama.cpp Vulkan DLLs** and compiles `dist\janus.exe`.

### 2. Download a model

Put a `.gguf` file in the `models\` folder. Use the included downloader:

```
go build -o dist\modelget.exe .\cmd\modelget
.\dist\modelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .\models\
```

Or download any GGUF from [Hugging Face](https://huggingface.co/models?library=gguf).

### 3. Configure

```
copy .env.example .env
```

Edit `.env`:

```
INFERENCE_BACKEND=vulkan
JANUS_MODEL_PATH=./models/Llama-3.2-3B-Instruct-Q8_0.gguf
JANUS_MAX_TOKENS=4096
```

| Variable | Default | Meaning |
| --- | --- | --- |
| `INFERENCE_BACKEND` | `vulkan` | `vulkan`, `cpu`, or `openrouter` |
| `JANUS_MODEL_PATH` | *(required)* | Path to your `.gguf` file |
| `JANUS_GPU_LAYERS` | `-1` | `-1` = all layers on GPU, `0` = CPU only |
| `JANUS_VRAM_CEILING_MB` | `9216` | VRAM budget hint (MiB) |
| `JANUS_MAX_TOKENS` | `4096` | Max tokens per reply |
| `JANUS_LISTEN_ADDR` | `127.0.0.1:8990` | Bind address |

### 4. Run

```
.\dist\janus.exe
```

Opens **<http://127.0.0.1:8990>** in your browser.

### 5. Verify

```
curl http://127.0.0.1:8990/health
```

---

## Quick start (Linux / macOS)

```
git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
go mod tidy
go build -o dist/janus ./cmd/janus
cp .env.example .env
# edit .env โ€” set JANUS_MODEL_PATH and INFERENCE_BACKEND=cpu if no Vulkan
./dist/janus
```

On Linux you need `libllama.so` next to the binary or on `LD_LIBRARY_PATH`.

---

## OpenAI-compatible API

```
curl http://127.0.0.1:8990/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
```

**Base URL:** `http://127.0.0.1:8990/v1`

### Endpoints

| Method | Path | Description |
| --- | --- | --- |
| GET | `/health` | Liveness check (`?deep=true` for details) |
| GET | `/v1/models` | Model list |
| POST | `/v1/chat/completions` | Chat (streaming supported) |
| POST | `/models/load` | Hot-swap model |
| GET | `/models/list` | Available .gguf files |
| GET | `/engine/status` | VRAM and backend info |

### Streaming

```
curl http://127.0.0.1:8990/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","stream":true,"messages":[{"role":"user","content":"Tell me a joke"}]}'
```

---

## Connect to Cursor / Cline / other clients

```
Base URL:  http://127.0.0.1:8990/v1
API Key:   (leave blank)
```

---

## Pitfalls (learned the hard way)

| Problem | What's going on | Fix |
| --- | --- | --- |
| **"It built but my changes aren't there"** | On Windows, Go can't overwrite a running `.exe`. | Stop all `janus.exe` in Task Manager, then rebuild. |
| **"Address already in use"** | A leftover process holds port 8990. | Task Manager โ†’ end all `janus.exe`. |
| **Server starts, chat fails** | `JANUS_MODEL_PATH` wrong or no `.gguf` in `models/`. | Set the path in `.env`, put the file in `models/`. |
| **"local engine failed to start"** | Missing `llama.dll` or GPU driver issue. | Run `.\build.ps1`. Update GPU drivers, or set `INFERENCE_BACKEND=cpu`. |
| **Wrong URL** | Default is `http://127.0.0.1:8990`, not 8080. | Bookmark 8990. |
| **First reply takes forever** | Model loading into VRAM โ€” normal. | Wait 10โ€“60s; smaller quants (`Q4`) load faster. |



---

## Project layout

```
cmd/janus/          Main server (OpenAI-compatible API)
cmd/modelget/       Hugging Face model downloader
internal/engine/    llama.cpp Vulkan/CPU backend
internal/bridge/    DLL loader and FFI bindings
internal/singleton/ Single-instance guard
models/             Put .gguf files here (not committed)
dist/               janus.exe + llama.dll after build
```

---

## Development

```
Get-Process -Name "janus" -ErrorAction SilentlyContinue | Stop-Process -Force
go test ./...
go build -o dist\janus.exe .\cmd\janus
.\dist\janus.exe
```

Or just `.\run.ps1`. For a full build including llama DLLs, use `.\build.ps1`.

Contributions welcome โ€” see [`CONTRIBUTING.md`](https://github.com/Vibra-Ingenn/Janus/blob/main/CONTRIBUTING.md).

---

## License

MIT โ€” see [`LICENSE`](https://github.com/Vibra-Ingenn/Janus/blob/main/LICENSE).
Maverick6179918
๐ŸŸง echo.github โญREADME: "Janus is a single Go binary that runs .gguf models on your machine (GPU or CPU) and exposes an OpenAI-compatible API. No Python, noVibra-Ingenn (GitHub org; HN submitter Maverick617, commits authored as "Janus Cleanup" <[email protected]>, co-authored with Claude Opus 4.6)โ€”โ€”

Interpretation history

Decision trace