Free & open source · Windows 11 on Snapdragon

Your Snapdragon NPU,
as a local AI server.

Calcine is a desktop app for Qualcomm GenieX. Download language models, chat with them on the Hexagon NPU, and give every app on your PC a secure, OpenAI-compatible API. Nothing leaves your machine.

Windows 11 ARM64 · GenieX included

Calcine's chat page: a conversation with Qwen3-4B running on the NPU, with the inference settings panel open.
Hexagon NPU AI Hub models, pre-compiled with QAIRT
Adreno GPU Any GGUF model, through llama.cpp
Oryon CPU The fallback that always works

Everything GenieX does, in one window

From first download to production API

Calcine wraps every GenieX command in a friendly interface, then adds what a CLI can't: a live view of your hardware, a model browser that knows your chip, and an API you can safely hand to other apps.

Chat with every option

Thinking, temperature, power mode, compute unit: settings adapt to each model's runtime. Streams with tokens per second.

Images and voice

Attach pictures or dictate to vision models. Attachments stay in memory, out of your saved conversations.

Find models that fit

Browse Qualcomm AI Hub for your chipset and search Hugging Face, with precisions, sizes, memory and disk checks.

See the NPU work

Live NPU, GPU, CPU and memory load, drivers, and a self-test that compares speed on each unit.

OpenAI-compatible API

On 127.0.0.1:18181, GenieX's default port. Per-app keys, a request queue and a request log.

Always up to date

GenieX ships in the installer and updates, or rolls back, from the app. Calcine updates itself with signed packages.

Discover

Pick the right precision, the first time

Calcine lists AI Hub models compiled for your exact chipset and searches Hugging Face for GGUF models. For each precision it shows the size and whether it fits in free memory and on disk, and recommends the one that runs best on the NPU.

  • Paste any Hugging Face, ModelScope or Docker Hub link
  • Downloads run in the background, with progress and cancel
  • Override the model type when a hub mislabels it
Choosing a precision for Qwen3-4B-GGUF: Q4_0 recommended, each option with its size and a memory check.

Hardware

Watch your NPU light up

Live load for the Hexagon NPU, Adreno GPU and Oryon CPU, read from Windows performance counters while a model generates. Run a self-test to see which unit is fastest for a given model.

  • Chipset detected automatically
  • Memory and model storage at a glance
  • GenieX, QAIRT and llama.cpp versions
The Hardware page with live NPU, GPU and CPU load charts, memory and storage bars.

Library

Your models, ready to run

Every installed model with its runtime, precisions and size. Switch a model between text and vision, copy its API id, or open a chat in one click.

  • Remove one precision or the whole model
  • Import GGUF folders or AI Hub bundles by drag and drop
  • Command palette for everything (Ctrl K)
The Library page listing three models with their runtime, precisions and size.

Local API

Point any OpenAI client at your NPU

Create a key per app in Server › API keys, then use the OpenAI SDK you already know. Streaming, vision and reasoning work as they do in the cloud, except everything runs on this PC.

  • Existing GenieX clients work unchanged
  • Requests queue while a model is busy
  • The log shows who called which model, never the prompts
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:18181/v1",
    api_key="calcine_…",  # Server › API keys
)
stream = client.chat.completions.create(
    model="qualcomm/Qwen3-4B",
    messages=[{"role": "user", "content": "Hello from my NPU!"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
The Server page with the API base URL, ready-to-copy client code and API key settings.

Secure by default

A local API that doesn't trust the browser

geniex serve has no authentication and accepts any web origin. Calcine never exposes it: it runs on a hidden loopback port, and every request goes through a gateway that checks, in order:

  1. 1

    Loopback only

    Listens on 127.0.0.1. Other machines on your network can't reach it.

  2. 2

    Host header

    Blocks DNS rebinding, where a web page points its domain at 127.0.0.1.

  3. 3

    Browser origin

    Web pages can't call the API from your browser unless you allow their origin.

  4. 4

    Per-app keys

    Shown once, stored as SHA-256 fingerprints, revocable instantly, scoped to inference or management.

  5. 5

    Request bodies

    Local file paths and internal URLs in requests are refused, so apps can't make GenieX read your files.

  6. 6

    Clean exit

    A Windows Job Object stops GenieX with Calcine, even after a crash. Nothing is left listening.

Signed updates, verified GenieX installers, and an SBOM with build provenance on every stable release. Read the security model

Get started in three steps

  1. Install Calcine. The installer also sets up GenieX, the Qualcomm runtime.
  2. Download a model. The welcome screen suggests a few that fit your PC.
  3. Chat, or connect an app. Create an API key and use any OpenAI client.

Windows 11 ARM64 · GenieX included

Calcine is in beta and not code-signed yet: Windows SmartScreen asks for confirmation before installing (More info › Run anyway).

Questions

Which PCs does Calcine run on?

Windows 11 on Arm with a Snapdragon processor that GenieX supports, such as the Snapdragon X Elite and X Plus. Linux on ARM64 is planned.

Does anything leave my PC?

No. Models run locally, the API only listens on this PC, and conversations are stored on this PC only. Calcine connects to the internet to download models and check for updates, and that's it: no telemetry.

Do I need to install GenieX first?

No. GenieX is bundled in the installer. You can update it, switch to pre-releases, or roll back from Hardware › GenieX runtime.

Which models can I run?

Qualcomm AI Hub models pre-compiled for the NPU, and GGUF models from Hugging Face run by llama.cpp on the NPU, GPU or CPU. Both text (LLM) and vision (VLM) models are supported.

I already use geniex serve. Will my apps still work?

Yes. Calcine listens on GenieX's default port, 18181, with the same OpenAI-compatible routes. Add an API key to your client, or turn off "Require an API key" for clients that can't send one.

Is Calcine made by Qualcomm?

No. Calcine is an independent open-source project under the MIT license. It redistributes GenieX under its BSD 3-Clause license.