2026-10-11 16:34 UTC

Quixotic AI claims its released Jinfer stack runs quantized chat, vision, audio, embedding, and speech models inside JVM applications without Python or sidecar services, potentially simplifying embedded local-inference deployment.

state: watchingheat: lowuncertainty: mediumnovelscott: lowinference-runtimes local-inference jvm-aiQuixotic AImukel

What is this?

Quixotic AI (the evolution of the llama4j project, driven by developer mukel) publishes an open-source, modular AI stack for the JVM headlined by jinfer, an in-process inference engine claiming chat, vision, audio, embeddings, reranking and text-to-speech with no Python, ONNX bridges, or sidecar services. Supporting modules include jota (a multi-backend tensor engine targeting Panama, C, CUDA, HIP, Metal, OpenCL and Mojo), jam (native quantized matmul kernels), pure-Java GGUF/Safetensors readers, and TikToken-compatible tokenizers, with first-class GraalVM Native Image support and Spring AI/LangChain4j integrations. The web results are almost entirely first-party (the project's own GitHub org and site); the launch's Reddit/HN posts and the Parakeet.java ASR port come from the same author, and nothing in the supplied material establishes independent benchmarking, production adoption, or third-party reproduction of the performance claims.

Why it matters to Scott

Still an example of a pattern Scott already runs rather than news about it: in-process, sidecar-free inference is a different deployment boundary from his Ollama-on-gamepc server setup, and his canon takes no position on JVM runtimes that Jinfer confirms or challenges. All evidence remains first-party (same author's launch posts and Parakeet port), engagement has fully decayed, and there is no active JVM project or independent validation that would make Scott re-evaluate his local-inference stack.
dev:technology.ollamadev:project.gamepcradar:concept.local-inferenceradar:concept.inference-runtimesradar:crispasr-local-audio-runtimeradar:rembed-pure-go-embeddings
queries asked of Scott's wikis
  • local inference stack choices — Ollama, llama.cpp, server vs in-process deployment boundary
  • embedded inference / sidecar-free model serving patterns in non-Python runtimes
  • JVM or Java projects Scott has built or evaluated; LangChain4j / Spring AI touchpoints
  • GGUF, quantization, and llama.cpp-format tooling Scott works with
  • RAG pipeline components — embeddings, reranking, ASR — and how he runs them locally
  • GraalVM native-image or single-binary distribution patterns in his tooling

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 633h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 07:26 (minted)⭐ origin echo-reconstructedPresents Jinfer as “AI in a jar,” with runnable JVM inference examples, Spring AI and LangChain4j integrations, and no sidecars, services, o
Quixotic AI on blog (echo) · attributed from hn.story.49708628 · published time unknown
—
09-15 06:42first on hacker news · published · lag ?Show HN: Jinfer – AI inference engine for the JVM. AI in a jar
mukel
—
09-15 08:09first on r/LocalLLaMA · published · lag ?jinfer: An open-source AI inference engine for the JVM. Finally, AI in jar.
mukel90
—
09-15 06:42amplified on hacker newshn.story.49708628
mukel
peak 6 · 1 comments · 9% of case engagement
09-15 08:09amplified on r/LocalLLaMA 👑reddit.post.1wgu62z
mukel90
peak 65 · 42 comments · 78% of case engagement
09-23 14:50amplified on r/LocalLLaMAreddit.post.1wo84pa
mukel90
peak 3 · 5 comments · 6% of case engagement
09-23 15:12amplified on hacker newshn.story.49817389
mukel
peak 3 · 1 comments · 5% of case engagement
09-24 18:19amplified on hacker newshn.story.49834739
mfiguiere
peak 1 · 0 comments · 1% of case engagement
09-15 07:21our radar first saw it · lag ?discovery anchor: hn.story.49708628—
pace: p70 vs 1032 stories at the 336h mark (now 633h old) — ahead of benzi-structure-resolved-code-intelligence (1.0x), behind perplexity-lily-apple-silicon-inference (1.0x)

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Jinfer – AI inference engine for the JVM. AI in a jar
Retrieved article excerpt

Open article · Retrieved 2026-09-15T07:22:27.579394+00:00

[Quixotic AI](https://qxotic.ai/)



# AI sovereignty for the JVM

Some call AI on the JVM quixotic.
  
We call it Quixotic AI.
  
**Finally, AI in a jar.**

[View on GitHub](https://github.com/qxoticai/qxotic)
[See it in action](https://qxotic.ai/#examples)

## Modular AI building blocks

Every component of the AI stack, a la carte.

[jinfer](https://github.com/qxoticai/qxotic/tree/main/jinfer)

AI inference engine. Chat, vision, embeddings and text-to-speech capabilities for the JVM, with Spring AI and LangChain4j integrations.

[jam](https://github.com/qxoticai/qxotic/tree/main/jam)

Fast quantized matrix-multiplication routines.

[Tok'n'Roll (toknroll)](https://github.com/qxoticai/qxotic/tree/main/toknroll)

TikToken-compatible and customizable tokenizers for popular LLM models.

[jota](https://github.com/qxoticai/qxotic/tree/main/jota)

Multi-backend tensor engine, in the works. Java, C, CUDA, HIP, Metal, OpenCL, and Mojo. **Write once, accelerate everywhere.**

[gguf](https://github.com/qxoticai/qxotic/tree/main/gguf)

Pure Java, read and write support for llama.cpp's GGUF model format.

[safetensors](https://github.com/qxoticai/qxotic/tree/main/safetensors)

Pure Java, read and write support for HuggingFace's Safetensors format.

## Examples

Runnable snippets using [jbang](https://www.jbang.dev).

Chat.java
TextToSpeech.java
Audio.java
Vision.java
Embed.java

Chat.java: LLM inference on the JVM
copy

```
jbang header + imports///usr/bin/env jbang "$0" "$@" ; exit $?
//JAVA 25+
//RUNTIME_OPTIONS --add-modules jdk.incubator.vector --enable-native-access=ALL-UNNAMED
//DEPS com.qxotic:jinfer-bom:0.2.0@pom
//DEPS com.qxotic:jinfer-spring-ai com.qxotic:jinfer-models-all
//DEPS com.qxotic:jam-native com.qxotic:jam-vector
//DEPS org.slf4j:slf4j-nop:2.0.17

import com.qxotic.jinfer.spring.ai.*;

void main() {
  try (var model = JinferChatModel.builder()
      .model("LiquidAI/LFM2.5-350M-GGUF:Q8_0")
      .build()) {
    System.out.println(model.call("What is the capital of France?"));
  }
}
```

$ jbang Chat.java  # "The capital of France is Paris."

TextToSpeech.java: Kokoro TTS on the JVM
copy

```
jbang header + imports///usr/bin/env jbang "$0" "$@" ; exit $?
//JAVA 25+
//RUNTIME_OPTIONS --add-modules jdk.incubator.vector --enable-native-access=ALL-UNNAMED
//DEPS com.qxotic:jinfer-bom:0.2.0@pom
//DEPS com.qxotic:jinfer-spring-ai
//DEPS com.qxotic:jinfer-kokoro
//DEPS com.qxotic:jam-native com.qxotic:jam-vector
//DEPS org.slf4j:slf4j-nop:2.0.17

import com.qxotic.jinfer.spring.ai.JinferSpeechModel;
import java.nio.file.*;

void main() throws Exception {
  try (var tts = JinferSpeechModel.builder()
      .model("simonfxr/kokoro.cpp-GGUF:Q8_0")
      .companion("voice", "simonfxr/kokoro.cpp-GGUF/voices/kokoro-voice-af_heart.gguf")
      .build()) {
    var line = "Turns out, matrix multiplications can talk. "
             + "The JVM just crunched a few billion numbers to say this.";
    Files.write(Path.of("kokoro.wav"), tts.call(line));
  }
}
```

$ jbang TextToSpeech.java  # kokoro.wav · seven seconds, in a warm voice · Kokoro-82M

Audio.java: transcription with Gemma 4 E2B
copy

```
jbang header + imports///usr/bin/env jbang "$0" "$@" ; exit $?
//JAVA 25+
//RUNTIME_OPTIONS --add-modules jdk.incubator.vector --enable-native-access=ALL-UNNAMED
//DEPS com.qxotic:jinfer-bom:0.2.0@pom
//DEPS com.qxotic:jinfer-spring-ai com.qxotic:jinfer-models-all
//DEPS org.springframework.ai:spring-ai-client-chat:2.0.1
//DEPS com.qxotic:jam-native com.qxotic:jam-vector
//DEPS org.slf4j:slf4j-nop:2.0.17

import com.qxotic.jinfer.spring.ai.*;
import org.springframework.ai.chat.client.ChatClient;
import org.springframework.core.io.UrlResource;
import org.springframework.util.MimeType;

void main() throws Exception {
  try (var gemma = JinferChatModel.builder()
      .model("unsloth/gemma-4-E2B-it-GGUF:Q4_K_M")
      .companion("media", "unsloth/gemma-4-E2B-it-GGUF/mmproj-BF16.gguf")
      .build()) {

    // jfk.wav: JFK's own voice, inaugural address, January 20, 1961
    var speech = new UrlResource("https://qxotic.ai/snippets/jfk.wav");
    var transcript = ChatClient.create(gemma).prompt()
        .user(u -> u.text("Transcribe this recording.")
                    .media(MimeType.valueOf("audio/wav"), speech))
        .call().content();
    System.out.println(transcript);
  }
}
```

$ jbang Audio.java  # "ask not what your country can do for you ..." · every word heard

Vision.java: images from a URL
copy

```
jbang header + imports///usr/bin/env jbang "$0" "$@" ; exit $?
//JAVA 25+
//RUNTIME_OPTIONS --add-modules jdk.incubator.vector --enable-native-access=ALL-UNNAMED
//DEPS com.qxotic:jinfer-bom:0.2.0@pom
//DEPS com.qxotic:jinfer-spring-ai com.qxotic:jinfer-models-all
//DEPS org.springframework.ai:spring-ai-client-chat:2.0.1
//DEPS com.qxotic:jam-native com.qxotic:jam-vector
//DEPS org.slf4j:slf4j-nop:2.0.17

import com.qxotic.jinfer.spring.ai.*;
import org.springframework.ai.chat.client.ChatClient;
import org.springframework.core.io.UrlResource;
import org.springframework.util.MimeTypeUtils;

void main() throws Exception {
  try (var vision = JinferChatModel.builder()
      .model("LiquidAI/LFM2.5-VL-3B-GGUF:Q8_0")
      .companion("media", "LiquidAI/LFM2.5-VL-3B-GGUF/mmproj-LFM2.5-VL-3B-Q8_0.gguf")
      .build()) {

    // The windmills of Consuegra, La Mancha: Don Quixote's "giants"
    var windmills = new UrlResource("https://qxotic.ai/snippets/consuegra.jpg");
    var answer = ChatClient.create(vision).prompt()
        .user(u -> u.text("Don Quixote saw giants here. What do you see?")
                    .media(MimeTypeUtils.IMAGE_JPEG, windmills))
        .call().content();
    System.out.println(answer);
  }
}
```

$ jbang Vision.java  # "several traditional windmills standing on a hill under a clear blue sky"

Embed.java: vector embeddings for RAG
copy

```
jbang header + imports///usr/bin/env jbang "$0" "$@" ; exit $?
//JAVA 25+
//RUNTIME_OPTIONS --add-modules jdk.incubator.vector --enable-native-access=ALL-UNNAMED
//DEPS com.qxotic:jinfer-bom:0.2.0@pom
//DEPS com.qxotic:jinfer-spring-ai com.qxotic:jinfer-models-all
//DEPS com.qxotic:jam-native com.qxotic:jam-vector
//DEPS org.slf4j:slf4j-nop:2.0.17

import com.qxotic.jinfer.spring.ai.JinferEmbeddingModel;

void main() {
  try (var emb = JinferEmbeddingModel.builder()
      .model("LiquidAI/LFM2.5-Embedding-350M-GGUF:Q8_0")
      .build()) {
    var a = emb.embed("AI on the JVM");
    var b = emb.embed("the JVM thinks now");
    var c = emb.embed("the python shed its skin");
    System.out.printf("similar   %.3f%n", cosineSimilarity(a, b));
    System.out.printf("unrelated %.3f%n", cosineSimilarity(a, c));
  }
}

cosineSimilarity helperfloat cosineSimilarity(float[] a, float[] b) {
  float dot = 0, na = 0, nb = 0;
  for (int i = 0; i < a.length; i++) {
    dot += a[i] * b[i]; na += a[i] * a[i]; nb += b[i] * b[i];
  }
  return (float) (dot / Math.sqrt(na * nb));
}
```

$ jbang Embed.java  # similar 0.609 · unrelated 0.081

## Benchmarks

Benchmarking is hard ... take these numbers as indicative, not definitive, always measure yourself.
  
`jinfer-bench` vs. `llama-bench` on CPU, ↑ higher is better.

## Built for the JVM from first principles

Local AI, end-to-end on the JVM: no sidecars, no services, no IPC.

### Just Java

No Python runtime, no ONNX, no glue code. Every component, from tokenizers to the inference engine, built for the JVM, not bolted onto it.

### GraalVM Native Image

Ships as a single self-contained binary: millisecond startup, small footprint, no JVM at runtime.

### Write Once, Accelerate Everywhere

One Tensor API, seven backends: Java, C, CUDA, HIP, Metal, OpenCL, and Mojo.
mukel61
🟧 echo.blog ⭐Presents Jinfer as “AI in a jar,” with runnable JVM inference examples, Spring AI and LangChain4j integrations, and no sidecars, services, oQuixotic AI——
🟠 redditjinfer: An open-source AI inference engine for the JVM. Finally, AI in jar.
LocalLLaMA
mukel906442
🟠 redditParakeet.java: Automatic Speech Recognition in pure Java.
LocalLLaMA
mukel9035
🟧 hnParakeet.java: Automatic Speech Recognition in Pure Javamukel31
🟧 hnJinfer-Parakeet: Automatic Speech Recognition for the JVMmfiguiere10

Interpretation history

Decision trace