An internal AI layer for your org, running on Claude and your own metal
inferenceEnJine turns your company's documents, tools, and workflows into one internal system. Claude handles the reasoning. Open models handle the high-volume work on hardware you already own. Every answer cites your own sources.
Built for ops, finance, and support teams that process documents and repeat the same judgment calls every day.
You own the code, the weights, and the keys. API spend goes only where a frontier model earns its cost.
One deployment, four working layers
Each layer ships running on your infrastructure, with evals and a runbook. Nothing waits in a proposal deck.
Knowledge layer
Your contracts, wikis, tickets, and reports become a queryable index on infrastructure you control. Claude reads the retrieved context and answers with citations back to your sources. Accuracy is measured on questions your own team writes, before anyone relies on the system.
Automation layer
Workflows in n8n move work between the tools your team already runs: pulling records, applying rules, writing outputs, and alerting a named human when a check fails. Claude handles the steps that need reading and judgment. Deterministic gates, retries, and audit trails handle everything else.
Model layer
High-volume, well-understood tasks get open models fine-tuned on your data: LoRA and QLoRA on single consumer GPUs. Hard reasoning stays on Claude. An eval harness routes each task to the layer that handles it at the accuracy you require, so the routing is a measured decision.
Cost layer
The audit writes your cost math before the build: which tasks need a frontier model, and which run cheaper on quantized open models (int8, 4-bit) hosted on machines you already own, including Apple Silicon. The target is exact API spend on reasoning and self-hosted serving for the rest.
Built on Claude
What Claude does in every deployment, and what stays on your hardware.
- Reasoning engine in the knowledge layer Claude reads retrieved context, follows multi-step instructions, and answers with grounded citations from your index.
- Judgment steps in the automation layer Document parsing, request classification, and drafting updates between systems run as Claude API calls inside the workflow.
- Agent tooling for setup and QA The same Claude-powered agents that build a deployment also verify it: harness design, eval runs, and pre-launch checks.
What stays on your hardware
- Document ingestion and chunking
- Embeddings and the retrieval index
- Quantized open models (int8 / 4-bit) for high-volume tasks
- The serving layer, behind your network
Your documents stay inside your deployment. Raw data never leaves it.
How a deployment runs
Fixed-scope engagements of 4 to 8 weeks. Weekly checkpoints, and every checkpoint leaves a working increment behind.
Audit
A map of your documents, tools, and current spend, then a one-page plan: the architecture, the hardware, the cost math, and the routing between Claude and self-hosted models. You approve the math before any build starts.
Deploy
The system grows in code you can read while it is being written: pipelines, indexes, eval harnesses, deploy scripts.
- Weekly checkpoint with a running system
- Evals measured on your data, not toy data
- Stop at any checkpoint, keep everything so far
Handover
Documentation, runbooks, and training for whoever operates it.
- You own the code, the weights, and the keys
- No per-seat license, no lock-in
- Operating cost written down before you pay it
Remote by default. Timezone CAT (UTC+2): overlap with EU mornings and US East afternoons. On-site available for African and EU cities by arrangement.
Proof, not claims
Three tracks of prior work. The numbers are measured and the artifacts are publicly inspectable.
Native inference engines
Built from-scratch C + Metal inference engines on Apple Silicon, bypassing PyTorch entirely. On a 48 GB Mac, int8 quantization cut resident model memory from 36.4 GB to 25.9 GB, and 4-bit reached about 19 GB, moving decode speed from 23 to 37 tokens per second. Memory headroom becomes a deployment decision you can make with numbers.
Fine-tuning under constraint
Fine-tuned open vision and language models under hard memory limits: gradient checkpointing, mixed precision, and batching choices that keep training alive on a single consumer GPU. A BERT fine-tune ships as a public demo at 90.7% accuracy. The same discipline an internal model gets: measured accuracy, on memory you can afford.
Production automation
A 50-node workflow in production: financial reconciliation across an accounting system, a Postgres mirror, and a payments API, verified to within a cent every month. It runs on schedule, checks itself before writing, and emails a named human when it fails. Client-named details stay out of this page by design; references available on request.
Get in touch
The first conversation is free. Describe your stack and the process you want to run internally, and you get an honest read on whether this helps before anyone commits to anything.