The open-source Private AI Platform.
Takeoffs counted from 190 sheets of drawings. Pricing and forecast across four ERPs. Voice agents on your phone lines. Intake that reads every file. Search walled by matter and role. Actions that wait for approval. One platform runs all of it, with inference on your own GPU, inside a tenancy you own on Azure, AWS, Google Cloud, Northflank or your own hardware. Every line auditable. Data locality mechanically verifiable, not marketed.
- GPU per tenant
- 1× H100-class
- 80 GB · native FP8 · in your tenancy
- License
- Apache 2.0
- OSI-approved · patent grant
- Vendor egress
- 0 bytes
- provable with tcpdump
Built for the work your data can't leave the building for.
Voice on your phone lines, documents in your drop folders, reporting over your ERPs, search inside your files, actions in your systems. FlatClaw is one platform for all of it, and every family below is backed by a real engagement: a deployment, a demo or a signed proposal, anonymized.
Estimating and quoting from drawings, specs and history
Takeoffs counted from the drawings' own data, reconciled against the team's proposal, priced on the company's rules, and benchmarked against what past jobs actually cost. Estimators steer in plain language; the final number waits for a human.
- 190
- drawing sheets read on one bid
- 1,045
- units itemized, 36 equipment families
- 88 · 216
- chillers and CRAHs, matching the estimators' count
- 161 vs 280
- the spread surfaced for adjudication
- Data-center construction groupDrawing takeoff and ROM estimating for hyperscale data-center bids
- Multi-brand industrial manufacturerConsolidation, forecast and pricing intelligence across four ERPs
- Flooring and concrete contractorEstimation benchmarks from twelve years of costing sheets
- European logistics groupQuote control tower over forwarding systems
Voice agents on your own lines
Two-way phone agents that work a voicemail queue, take an intake call, or answer a front-line number, with the speech, the reasoning and the recordings all inside your tenancy.
Documents in, structured data out
Agents that watch drop folders and inboxes, read placement files, drawings and forms, map them onto your schema, and hand a clean file to the system that already runs the process.
Reporting across every system you run
Consolidated financials across four ERPs, a quoting control tower over forwarding systems, clinical reports over device data: governed connectors, one lakehouse, questions in plain English.
Knowledge search with walls in it
Retrieval scoped by matter, role and account before the model sees a document. A lawyer searches only their matters; a banker's assistant never sees account data it shouldn't.
Operations that pause for approval
Scheduled work, compliance gates and back-office automation where every consequential action waits for a human, replays with that person's own credentials, and lands in the audit trail.
Sales, marketing and content
CRM-connected agents with per-user credentials, contact data refreshed through a waterfall of sources, franchise coaching over live numbers, and on-brand proposals from a shared skill.
Estimating is where the platform earns its keep.
Bids are the most expensive documents a company produces and the most data-starved. FlatClaw reads the drawings, the specifications and the job history, counts what is actually on the page, prices it on your own rules, and shows every source, inside your tenancy, because a bid is the last thing that should leave the building.
Takeoff counted from the drawings themselves.
A national pre-construction group quotes hundreds of bids a month with a ten-person estimating team, and one hyperscale estimate can take a single estimator two months. What is on the printed page is what the subcontractor is on the hook for.
- Four issue-for-construction drawing sets, 190 sheets and three manual volumes read natively: 1,045 units across 36 equipment families, every count citing its sheets.
- Reconciled with the estimators' own proposal on the lines that drive price: 88 chillers, 216 computer-room air handlers, 36 pumps.
- A 161-versus-280 fan-wall-unit spread between drawings and proposal surfaced for adjudication instead of averaged away. That gap is the miss a sub otherwise eats.
- Estimators steer in plain language (increase all labor by five percent, union state, twenty percent spares); guardrails cap what moves without a manager; the estimate waits for approval.
- Rough-order-of-magnitude output in the team's own schedule-of-pricing format, then proposals in the house format.
Pricing, margin and forecast across four ERPs.
Five brands built by acquisition, four ERPs from four eras, one CRM, and a month-end that lived in spreadsheets. The controller wanted a ledger tape back to the transaction before trusting a consolidated number.
- A governed lakehouse in the company's own Azure tenant on Microsoft Fabric, fed from every ERP and the CRM, with lineage to the source transaction.
- Forecast and pipeline by business unit with accuracy tracked over time, large-job margin watch, and an AR and collections cockpit on the same store.
- Ask about any strategic account and get revenue, margin, pipeline and whitespace across brands, with the CRM written back.
- Pricing intelligence and aftermarket analytics next, on the same foundation, because the estimate and the actual finally live in one place.
- The next acquisition onboards as a templated pattern priced in weeks, not a project each time.
A platform you deploy, not a service you rent.
The frontier labs have shown what AI can do inside an organization: agents with a task inbox, scheduled work, document memory, direct access to files and connected apps. Claude Cowork, Gemini Enterprise Agent and GPT‑6 + Atlas all deliver it the same way: as a service, on their servers, with your data sent over on every request.
The problem
For a large part of the economy that is a non-starter. Law firms, healthcare, banks, collections, manufacturers with confidential financials, anyone with a contract that says the data stays put: the most capable AI on the market is the AI they are not allowed to use. The usual fallback is a narrow point solution per problem, each with its own vendor and its own copy of the data.
The answer
FlatClaw is the whole platform, deployed into a tenancy you own: inference on your own GPU, agents, voice, retrieval, connectors, approvals, scheduling and memory, with role-based access at every tool call and an audit trail under all of it. Build every use case on the same foundation instead of buying a vendor per problem. Open source, Apache 2.0, every line yours to read.
Runs on the cloud you already trust.
FlatClaw is a set of containers and one GPU node. It deploys into a tenancy the customer owns on any of these, with the same image, the same control plane, and the same privacy proof.
The lane for Microsoft-first organizations.
Your account, your VPC, no public inference endpoint.
A project the customer owns, fenced with VPC Service Controls.
The fastest path from zero to a running tenant.
For data that cannot be in any cloud at all.
Microsoft Azure, Amazon Web Services, Google Cloud and Northflank are trademarks of their respective owners. Logos identify supported deployment targets; no endorsement is implied.
Everything a Private AI Platform needs, pre-integrated.
Eight components, one image, one tenancy. Every use case above runs on the same eight. Each one is replaceable and auditable on its own.
Next.js 16 + React 19 product surface. Chat, agent fleet, approvals, cron, MCP services, workspace files, Memory, Admin (RBAC). SSE-streamed tool use.
Built on the minimal open Pi agent core. Sessions, multi-step planning, sandboxed tool execution, RBAC enforced at every tool call, scheduling and approvals above it. Owns per-agent memory.
Patched SGLang + Gemma 4 31B Dense on a single NVIDIA H100-class GPU (80 GB, native FP8) inside the customer's own tenancy, served at the model's native 256K context.
The harness's built-in per-agent SQLite memory — keyword (BM25) search over each agent's MEMORY.md and memory/ files, seeded automatically for every agent. The agent maintains it across sessions. Semantic recall via bge-m3 lands in v0.4.
First-party Model Context Protocol servers: Google (Gmail/Calendar/Drive/Docs/Sheets) and Jira, plus private add-on connectors through the same plugin registry. Consequential actions pause for human approval and replay with the user's own credentials. Per-user credentials scoped (tenant, user, service), never tenant-wide.
Multiple users per tenant. Per-user Tool Access (allow/deny over built-in + MCP tools) on the harness's native tool policy, plus always-on cross-user isolation. Per-user credentials scoped (tenant, user, service).
Each customer gets their own tenancy — an Azure resource group, an AWS account, a Google Cloud project, a Northflank project, or an on-prem cluster. Strict isolation, dedicated GPU, no shared state across tenants.
ghcr.io/skytruax/flatclaw-inference:latest — public on GHCR, ~18 GB, no baked weights. Every deployment pulls the same image.
Single-tenant. Customer-owned. End-to-end.
Everything — Portal, the agent harness, Inference (GPU), and the weights-server — lives in one tenancy on the cloud the customer already runs: an Azure resource group, an AWS account, a Google Cloud project, a Northflank project, or a rack in their building. The customer holds the account. Nothing leaves it.
tools.deny — always-on cross-user roster isolation (each agent sees only its own servers' tools) plus a per-user Tool Access panel that toggles built-in and MCP tools off; denied tools are filtered from the roster before the model sees them.MEMORY.md + memory/ files. Semantic recall via bge-m3 (on its own GPU card) and RAGFlow cited-document retrieval land in v0.4.weights-server pod over a tenancy-local volume — staged once, never moved at boot.≈ $2,000 / month per tenant. Every use case, one flat rate.
One GPU carries a tenant's whole workload: voice, intake, reporting, search and agents share it. Indicative monthly cost for a single tenant held warm 24/7 on a managed H100 plan, at the reference lane's published list pricing. Azure and AWS H100 classes land in the same band on reserved terms; bare metal amortizes lower. The H100 dominates; everything else combined is under $200. The rate is per tenant and scales with the tenant — not metered per token or per seat.
Monthly cost breakdown
| Inference (H100 80GB, held warm) | ~$1,800 |
| Portal — small compute (4 vCPU / 8 GB) | ~$50 |
| Agent harness — small compute | ~$50 |
| RAGFlow + corpus volume | ~$30 |
| weights-server + 200 GB nvme | ~$30 |
| Egress · TLS · observability | included |
| Total per tenant, all-in | ~$2,000 / mo |
List prices, round numbers. Committed-use or annual terms on any of the clouds typically reduce the GPU line. One bill, from the cloud the customer already has a relationship with.
How one H100 carries a tenant — and how it scales
What the GPU serves is peak concurrent active sessions, not the tenant's total user count — people skim a result, edit a doc, take a call, ask a follow-up. The H100 is sized to that concurrent peak; the per-tenant rate doesn't move with seat count.
One H100 SGLang process at Gemma 4 31B FP8 sustains ~8–12 concurrent streaming chats with first-token latency in the 1–2 s range. SGLang's RadixAttention prefix cache earns most of that on conversational reuse.
Memory recall, RAG retrieval, file reads, OAuth tool invocations — all gateway- or skill-side. The LLM is invoked for chat turns and tool-call planning. A typical session is a handful of LLM calls, not hundreds.
Gemma 4 31B FP8 (~33 GB) + KV cache + bge-m3 fits in 80 GB with ~25 GB free. The v0.4 cascade lands a co-resident smaller Gemma in that headroom for fast-turn / planning traffic — same hardware, ~2× concurrent capacity.
When a tenant outgrows one card, the next step is a higher-tier GPU plan or a multi-GPU node on the same cloud — or a second inference service for triage. Same tenancy, same architecture, same per-tenant model.
Mechanically provable, not marketed.
The privacy story is not a marketing claim. It is a test you can run yourself, on whichever cloud you deploy to.
- 1Provision a tenant in your own cloud tenancy — Azure, AWS, Google Cloud, Northflank, or your own hardware.
- 2Exercise the shipped features end-to-end (chat, memory, MCP services with approval-gated actions, scheduled-task fire, GPU cold-boot). As features land, each is added to this test loop.
- 3Run tcpdump on the tenancy's egress for the full session.
- 4Confirm zero packets to Anthropic, OpenAI, Google AI, Hugging Face, ElevenLabs, Chroma Cloud, or any third-party inference endpoint. Inference traffic stays inside the tenancy — Portal → Gateway → GPU is all internal network. The only external egress: services the user explicitly connected via OAuth.
This check runs mechanically on every release. It is the promise the project exists to keep.
Best-in-class open-source, end to end.
Every dependency is MIT / Apache / BSD compatible. Nothing here is a vendor lock-in — including the cloud.
Shipping in the open.
What's working today vs. what's coming next is honest, enumerated, and verifiable.
v0.3.0
This release- ▸Human approval engine — consequential MCP actions (outbound mail, destructive/exposing calls) are composed, pause for human sign-off in the Portal, and replay with the user's own credentials on approve
- ▸Public/private MCP split — the repo ships mcp/public (Google, Jira); private add-on connectors self-register through the same plugin registry
- ▸Harness runtime pin bumped and re-verified against the RBAC contract (tool-name flattening, deny-glob pipeline)
- ▸Portal approvals queue + per-service admin visibility controls
v0.4
Next- ▸One-command tenant provisioning — provision-tenant.sh / destroy-tenant.sh: full tenant lifecycle on the target cloud (Northflank lane first; Azure, AWS and Google Cloud lanes follow)
- ▸RAGFlow — cited document retrieval behind a stable interface
- ▸Semantic memory + embeddings via bge-m3 — on its own GPU card
- ▸Scrapling web fetch; additional CRM/ERP connectors as add-on services
- ▸Voice — VoxCPM2 open-weight cloning + TTS; Image — ComfyUI + SDXL
- ▸Cascade routing + TurboQuant turbo4 — 1M-token context on a single card
v0.5+
Future- ▸WorkOS SSO for enterprise tenants (Okta / Azure AD / Google Workspace)
- ▸Optional shared-GPU multi-tenancy for an entry tier below the dedicated-GPU threshold
- ▸Audio/video transcription ingest in RAGFlow
- ▸A "studio" for users to author their own skills
Pull it. Audit it. Run it.
Apache 2.0 with an explicit patent grant. Bring your own cloud, or your own hardware. Or bring a workload and see the platform run it.