FlatClaw, Private AI Platform
v0.3.0 · Apache 2.0

The open-source Private AI Platform.

Takeoffs counted from 190 sheets of drawings. Pricing and forecast across four ERPs. Voice agents on your phone lines. Intake that reads every file. Search walled by matter and role. Actions that wait for approval. One platform runs all of it, with inference on your own GPU, inside a tenancy you own on Azure, AWS, Google Cloud, Northflank or your own hardware. Every line auditable. Data locality mechanically verifiable, not marketed.

GPU per tenant
1× H100-class
80 GB · native FP8 · in your tenancy
License
Apache 2.0
OSI-approved · patent grant
Vendor egress
0 bytes
provable with tcpdump
One platform, every kind of work

Built for the work your data can't leave the building for.

Voice on your phone lines, documents in your drop folders, reporting over your ERPs, search inside your files, actions in your systems. FlatClaw is one platform for all of it, and every family below is backed by a real engagement: a deployment, a demo or a signed proposal, anonymized.

Featured family

Estimating and quoting from drawings, specs and history

Takeoffs counted from the drawings' own data, reconciled against the team's proposal, priced on the company's rules, and benchmarked against what past jobs actually cost. Estimators steer in plain language; the final number waits for a human.

190
drawing sheets read on one bid
1,045
units itemized, 36 equipment families
88 · 216
chillers and CRAHs, matching the estimators' count
161 vs 280
the spread surfaced for adjudication

Voice agents on your own lines

Two-way phone agents that work a voicemail queue, take an intake call, or answer a front-line number, with the speech, the reasoning and the recordings all inside your tenancy.

Documents in, structured data out

Agents that watch drop folders and inboxes, read placement files, drawings and forms, map them onto your schema, and hand a clean file to the system that already runs the process.

Reporting across every system you run

Consolidated financials across four ERPs, a quoting control tower over forwarding systems, clinical reports over device data: governed connectors, one lakehouse, questions in plain English.

Knowledge search with walls in it

Retrieval scoped by matter, role and account before the model sees a document. A lawyer searches only their matters; a banker's assistant never sees account data it shouldn't.

Operations that pause for approval

Scheduled work, compliance gates and back-office automation where every consequential action waits for a human, replays with that person's own credentials, and lands in the audit trail.

Sales, marketing and content

CRM-connected agents with per-user credentials, contact data refreshed through a waterfall of sources, franchise coaching over live numbers, and on-brand proposals from a shared skill.

Estimating, in depth

Estimating is where the platform earns its keep.

Bids are the most expensive documents a company produces and the most data-starved. FlatClaw reads the drawings, the specifications and the job history, counts what is actually on the page, prices it on your own rules, and shows every source, inside your tenancy, because a bid is the last thing that should leave the building.

Hyperscale data-center bids

Takeoff counted from the drawings themselves.

A national pre-construction group quotes hundreds of bids a month with a ten-person estimating team, and one hyperscale estimate can take a single estimator two months. What is on the printed page is what the subcontractor is on the hook for.

  • Four issue-for-construction drawing sets, 190 sheets and three manual volumes read natively: 1,045 units across 36 equipment families, every count citing its sheets.
  • Reconciled with the estimators' own proposal on the lines that drive price: 88 chillers, 216 computer-room air handlers, 36 pumps.
  • A 161-versus-280 fan-wall-unit spread between drawings and proposal surfaced for adjudication instead of averaged away. That gap is the miss a sub otherwise eats.
  • Estimators steer in plain language (increase all labor by five percent, union state, twenty percent spares); guardrails cap what moves without a manager; the estimate waits for approval.
  • Rough-order-of-magnitude output in the team's own schedule-of-pricing format, then proposals in the house format.
Read the spotlight
Industrial manufacturer, five brands

Pricing, margin and forecast across four ERPs.

Five brands built by acquisition, four ERPs from four eras, one CRM, and a month-end that lived in spreadsheets. The controller wanted a ledger tape back to the transaction before trusting a consolidated number.

  • A governed lakehouse in the company's own Azure tenant on Microsoft Fabric, fed from every ERP and the CRM, with lineage to the source transaction.
  • Forecast and pipeline by business unit with accuracy tracked over time, large-job margin watch, and an AR and collections cockpit on the same store.
  • Ask about any strategic account and get revenue, margin, pipeline and whitespace across brands, with the CRM written back.
  • Pricing intelligence and aftermarket analytics next, on the same foundation, because the estimate and the actual finally live in one place.
  • The next acquisition onboards as a templated pattern priced in weeks, not a project each time.
Read the spotlight
Why it exists

A platform you deploy, not a service you rent.

The frontier labs have shown what AI can do inside an organization: agents with a task inbox, scheduled work, document memory, direct access to files and connected apps. Claude Cowork, Gemini Enterprise Agent and GPT‑6 + Atlas all deliver it the same way: as a service, on their servers, with your data sent over on every request.

The problem

For a large part of the economy that is a non-starter. Law firms, healthcare, banks, collections, manufacturers with confidential financials, anyone with a contract that says the data stays put: the most capable AI on the market is the AI they are not allowed to use. The usual fallback is a narrow point solution per problem, each with its own vendor and its own copy of the data.

The answer

FlatClaw is the whole platform, deployed into a tenancy you own: inference on your own GPU, agents, voice, retrieval, connectors, approvals, scheduling and memory, with role-based access at every tool call and an audit trail under all of it. Build every use case on the same foundation instead of buying a vendor per problem. Open source, Apache 2.0, every line yours to read.

Cloud partners

Runs on the cloud you already trust.

FlatClaw is a set of containers and one GPU node. It deploys into a tenancy the customer owns on any of these, with the same image, the same control plane, and the same privacy proof.

Microsoft AzureMicrosoft Azure
Delivered with an implementation partner

The lane for Microsoft-first organizations.

Amazon Web Services
Delivered with an implementation partner

Your account, your VPC, no public inference endpoint.

Google Cloud
Delivered with an implementation partner

A project the customer owns, fenced with VPC Service Controls.

NorthflankNorthflank
Reference lane · scripted today

The fastest path from zero to a running tenant.

Your own hardware
Bring your own metal

For data that cannot be in any cloud at all.

How each lane works

Microsoft Azure, Amazon Web Services, Google Cloud and Northflank are trademarks of their respective owners. Logos identify supported deployment targets; no endorsement is implied.

What's in the box

Everything a Private AI Platform needs, pre-integrated.

Eight components, one image, one tenancy. Every use case above runs on the same eight. Each one is replaceable and auditable on its own.

FlatClaw Portal

Next.js 16 + React 19 product surface. Chat, agent fleet, approvals, cron, MCP services, workspace files, Memory, Admin (RBAC). SSE-streamed tool use.

Agent harness

Built on the minimal open Pi agent core. Sessions, multi-step planning, sandboxed tool execution, RBAC enforced at every tool call, scheduling and approvals above it. Owns per-agent memory.

Inference

Patched SGLang + Gemma 4 31B Dense on a single NVIDIA H100-class GPU (80 GB, native FP8) inside the customer's own tenancy, served at the model's native 256K context.

Per-agent memory

The harness's built-in per-agent SQLite memory — keyword (BM25) search over each agent's MEMORY.md and memory/ files, seeded automatically for every agent. The agent maintains it across sessions. Semantic recall via bge-m3 lands in v0.4.

MCP services

First-party Model Context Protocol servers: Google (Gmail/Calendar/Drive/Docs/Sheets) and Jira, plus private add-on connectors through the same plugin registry. Consequential actions pause for human approval and replay with the user's own credentials. Per-user credentials scoped (tenant, user, service), never tenant-wide.

RBAC + per-user creds

Multiple users per tenant. Per-user Tool Access (allow/deny over built-in + MCP tools) on the harness's native tool policy, plus always-on cross-user isolation. Per-user credentials scoped (tenant, user, service).

Single-tenant by design

Each customer gets their own tenancy — an Azure resource group, an AWS account, a Google Cloud project, a Northflank project, or an on-prem cluster. Strict isolation, dedicated GPU, no shared state across tenants.

One image, every tenant

ghcr.io/skytruax/flatclaw-inference:latest — public on GHCR, ~18 GB, no baked weights. Every deployment pulls the same image.

Architecture

Single-tenant. Customer-owned. End-to-end.

Everything — Portal, the agent harness, Inference (GPU), and the weights-server — lives in one tenancy on the cloud the customer already runs: an Azure resource group, an AWS account, a Google Cloud project, a Northflank project, or a rack in their building. The customer holds the account. Nothing leaves it.

Customer's cloud tenancy · one per tenant · Azure / AWS / Google Cloud / Northflank / on-prem
Browser
The user — an admin or end-user inside the customer's org
HTTPS · cookie-auth
FlatClaw Portal
Next.js 16 + React 19 + SQLite
ChatAgentsApprovalsCronMCP servicesMemoryAdmin
server-owned WebSocket · ws://:18789
Agent harness
Built on the minimal open Pi agent core · sessions · tool dispatch · RBAC enforced at every tool call
MCP (per-agent, deny-glob scoped) · per-user Tool Access (native tools.deny)
Google
Gmail · Calendar · Drive · Docs · Sheets · Contacts
Add-on connectors
CRM · ERP · private services
Jira
Atlassian Cloud
Sandbox
per-tool exec · per-user scoped credentials
MCP services are first-party servers the agent calls over Model Context Protocol; per-user credentials scoped per (tenant, user, service). RBAC is the harness's native per-agent tools.deny — always-on cross-user roster isolation (each agent sees only its own servers' tools) plus a per-user Tool Access panel that toggles built-in and MCP tools off; denied tools are filtered from the roster before the model sees them.
Per-agent memory uses the harness's built-in per-agent SQLite engine — keyword (BM25) search over each agent's MEMORY.md + memory/ files. Semantic recall via bge-m3 (on its own GPU card) and RAGFlow cited-document retrieval land in v0.4.
internal tenancy network · TLS · bearer-authenticated
Inference (GPU) · same tenancy
Inference service
Patched SGLang · Gemma 4 31B-IT (FP8) · 256K context · NVIDIA H100 (80 GB · sm_90 · native FP8)
Weights served by the in-project weights-server pod over a tenancy-local volume — staged once, never moved at boot.
No vendor egress
Zero packets to Anthropic, OpenAI, Google AI, Hugging Face, ElevenLabs, or any third-party inference endpoint. Verifiable with tcpdump.
Customer holds the account
Customer's own cloud account — Azure, AWS, Google Cloud, Northflank, or their own hardware — billed directly to them. We never touch the bill or the data.
One image, every tenant
ghcr.io/skytruax/flatclaw-inference:latest. SGLang base + entrypoint, no baked weights. Public, auditable, reproducible.
Token Economics

≈ $2,000 / month per tenant. Every use case, one flat rate.

One GPU carries a tenant's whole workload: voice, intake, reporting, search and agents share it. Indicative monthly cost for a single tenant held warm 24/7 on a managed H100 plan, at the reference lane's published list pricing. Azure and AWS H100 classes land in the same band on reserved terms; bare metal amortizes lower. The H100 dominates; everything else combined is under $200. The rate is per tenant and scales with the tenant — not metered per token or per seat.

Monthly cost breakdown

Inference (H100 80GB, held warm)~$1,800
Portal — small compute (4 vCPU / 8 GB)~$50
Agent harness — small compute~$50
RAGFlow + corpus volume~$30
weights-server + 200 GB nvme~$30
Egress · TLS · observabilityincluded
Total per tenant, all-in~$2,000 / mo

List prices, round numbers. Committed-use or annual terms on any of the clouds typically reduce the GPU line. One bill, from the cloud the customer already has a relationship with.

How one H100 carries a tenant — and how it scales

Concurrency, not headcount, sets the load.

What the GPU serves is peak concurrent active sessions, not the tenant's total user count — people skim a result, edit a doc, take a call, ask a follow-up. The H100 is sized to that concurrent peak; the per-tenant rate doesn't move with seat count.

The 31B path handles 8–12 concurrent streams.

One H100 SGLang process at Gemma 4 31B FP8 sustains ~8–12 concurrent streaming chats with first-token latency in the 1–2 s range. SGLang's RadixAttention prefix cache earns most of that on conversational reuse.

Most user actions don't touch the LLM at all.

Memory recall, RAG retrieval, file reads, OAuth tool invocations — all gateway- or skill-side. The LLM is invoked for chat turns and tool-call planning. A typical session is a handful of LLM calls, not hundreds.

Headroom for bursts, then a cascade.

Gemma 4 31B FP8 (~33 GB) + KV cache + bge-m3 fits in 80 GB with ~25 GB free. The v0.4 cascade lands a co-resident smaller Gemma in that headroom for fast-turn / planning traffic — same hardware, ~2× concurrent capacity.

Tenants scale the GPU plan, not the architecture.

When a tenant outgrows one card, the next step is a higher-tier GPU plan or a multi-GPU node on the same cloud — or a second inference service for triage. Same tenancy, same architecture, same per-tenant model.

Private LLM

Mechanically provable, not marketed.

The privacy story is not a marketing claim. It is a test you can run yourself, on whichever cloud you deploy to.

  1. 1Provision a tenant in your own cloud tenancy — Azure, AWS, Google Cloud, Northflank, or your own hardware.
  2. 2Exercise the shipped features end-to-end (chat, memory, MCP services with approval-gated actions, scheduled-task fire, GPU cold-boot). As features land, each is added to this test loop.
  3. 3Run tcpdump on the tenancy's egress for the full session.
  4. 4Confirm zero packets to Anthropic, OpenAI, Google AI, Hugging Face, ElevenLabs, Chroma Cloud, or any third-party inference endpoint. Inference traffic stays inside the tenancy — Portal → Gateway → GPU is all internal network. The only external egress: services the user explicitly connected via OAuth.

This check runs mechanically on every release. It is the promise the project exists to keep.

Technology

Best-in-class open-source, end to end.

Every dependency is MIT / Apache / BSD compatible. Nothing here is a vendor lock-in — including the cloud.

InferencePatched SGLang + Gemma 4 31B Dense
SiliconNVIDIA H100-class · 80 GB · native FP8
SubstrateYour cloud — Azure, AWS, Google Cloud, Northflank, or bare metal — one tenancy per customer
ContextTurboQuant turbo4 KV — 1M tokens on a single card (roadmap)
Agent harnessBuilt on the minimal open Pi agent core — RBAC at every tool call · per-agent memory built in · gateway layer swappable behind the session API
FrontendNext.js 16 + React 19 + TypeScript + SQLite
Authbetter-auth (v1) · WorkOS SSO (v2)
MemoryHarness-native per-agent SQLite — keyword search, seeded per agent
RetrievalRAGFlow — cited document answers (v0.4)
Embeddings (v0.4)bge-m3 — semantic memory + RAG, on its own GPU card
Voice (v0.4)VoxCPM2 — open-weight cloning + TTS
Image (v0.4)ComfyUI + SDXL
Roadmap

Shipping in the open.

What's working today vs. what's coming next is honest, enumerated, and verifiable.

0.3.0

v0.3.0

This release
  • Human approval engine — consequential MCP actions (outbound mail, destructive/exposing calls) are composed, pause for human sign-off in the Portal, and replay with the user's own credentials on approve
  • Public/private MCP split — the repo ships mcp/public (Google, Jira); private add-on connectors self-register through the same plugin registry
  • Harness runtime pin bumped and re-verified against the RBAC contract (tool-name flattening, deny-glob pipeline)
  • Portal approvals queue + per-service admin visibility controls
0.4

v0.4

Next
  • One-command tenant provisioning — provision-tenant.sh / destroy-tenant.sh: full tenant lifecycle on the target cloud (Northflank lane first; Azure, AWS and Google Cloud lanes follow)
  • RAGFlow — cited document retrieval behind a stable interface
  • Semantic memory + embeddings via bge-m3 — on its own GPU card
  • Scrapling web fetch; additional CRM/ERP connectors as add-on services
  • Voice — VoxCPM2 open-weight cloning + TTS; Image — ComfyUI + SDXL
  • Cascade routing + TurboQuant turbo4 — 1M-token context on a single card
0.5+

v0.5+

Future
  • WorkOS SSO for enterprise tenants (Okta / Azure AD / Google Workspace)
  • Optional shared-GPU multi-tenancy for an entry tier below the dedicated-GPU threshold
  • Audio/video transcription ingest in RAGFlow
  • A "studio" for users to author their own skills
Get started

Pull it. Audit it. Run it.

Apache 2.0 with an explicit patent grant. Bring your own cloud, or your own hardware. Or bring a workload and see the platform run it.