AI Delivery••10 min read

How Metacto Gave Every Employee a Private AI Assistant in One Week

We ran OpenClaw, Hermes, Grok Bot, Instinct, Muse, and Dots on real work. Then we replaced them all with one open-source assistant inside our own cloud. The runtime took a day. The controls took the rest of the week.

Chris Fitkin
Chris Fitkin
Partner & Co-Founder

Personal AI assistants are all the rage with consumers. People use them to plan trips, draft emails, summarize documents, and think out loud. Then they come to work and the same assistant is either banned, ungoverned, or quietly pasted full of company data.

We had the same problem at Metacto. So last week we fixed it for ourselves.

Every US-based Metacto employee now has a private AI assistant. It runs in Metacto’s cloud, signs in with our identity provider, reaches only the systems each person is allowed to reach, uses models we approved, and reports every action and every dollar to a dashboard leadership can read.

This post is how we did it, and what we would tell any leadership team about to do the same.

The short version

The assistant runtime took a day. Identity, permissions, connectors, logging, model routing, and spend controls took the rest of the week. That ratio is the whole lesson.

How We Got Here: A Year of Running Every Assistant

We did not start with a plan. We started by running everything.

Over the past year, Metacto has used nearly every personal AI assistant worth trying: open-source agents we hosted ourselves, and the wave of hosted personal agents from the large AI companies. We ran them on real work, with real inboxes and calendars, because that is the only way to learn what holds up.

Stage 1: OpenClaw bots. We were early to OpenClaw, the self-hosted open-source assistant that became one of the most-starred projects on GitHub. We stood up OpenClaw bots for different people and teams, connected them to messaging apps and work tools, and they were genuinely useful within hours. They also multiplied. Each bot had its own host, its own credentials, its own prompts, and an owner who was busy with client work. Nobody could answer simple questions: who has access to what, which models are we paying for, and what did the bots do last week?

That is a small version of the problem we see at clients: AI sprawl. Many tools, no shared control plane.

Stage 2: Self-improving open-source agents. Next we ran Hermes Agent from Nous Research, which adds persistent memory and a learning loop that turns experience into reusable skills. It showed how quickly open-source runtimes were closing the gap with commercial products. It did not change the operating problem. A smarter agent with no governance around it is still ungoverned.

Stage 3: Hosted personal agents. Then the large AI companies shipped their own. We ran xAI’s Grok Bot, Instinct, Meta’s Muse, and OpenAI’s Dots. They are impressive. They keep working while you are away, act across your apps, and remember context. For personal life, they are a superpower.

For a company, every one of them asks for the same trade: connect your company inbox, calendar, documents, and systems to the vendor’s cloud, use the vendor’s models, and trust the vendor’s view of what happened. Multiply that by every employee and every new assistant that launches next month, and you have a data-governance problem with a subscription fee.

What a year of assistants taught us

The assistants kept getting better. The questions leadership asked about cost, access, training, and audit stayed unanswered.

What we ran: OpenClaw bots

What worked
Self-hosted, fast to stand up, huge integration ecosystem.
Why we moved on
Every bot was its own deployment. No shared identity, permissions, spend visibility, or audit.

What we ran: Hermes Agent

What worked
Open source, persistent memory, skills that improve with use.
Why we moved on
Better agent, same operating gap. Capability was never the missing piece.

What we ran: Grok Bot, Instinct, Muse, Dots

What worked
Polished, persistent agents that act on your behalf across apps.
Why we moved on
Company data and model choice live in a vendor's cloud, one vendor and one employee at a time.

The lesson. The assistant experience is becoming commodity infrastructure. Foundation-model companies and open-source communities ship improvements to chat, memory, tool use, and agent behavior every few weeks, and the leader changes every quarter. Metacto is not going to out-ship them at the interface layer, and neither is your internal team. What none of them gave us was control.

So we made a rule for ourselves:

Own the management layer, not the commodity layer.

We would adopt the best open-source assistant runtime available, run it inside our own environment, retire the OpenClaw bots, and put our engineering effort into everything around it.

How We Picked the Runtime

We evaluated the strongest permissively licensed (MIT and similar) assistant platforms against a short list of requirements. Feature depth mattered less than how well each one could be controlled.

Runtime evaluation criteria

We picked the runtime that best fit these criteria this quarter. We built everything else on the assumption that we will replace it.

Requirement: Self-hostable

What we checked
Runs entirely in our cloud account with no required calls to a vendor's hosted service.
Why it mattered
Company data and conversation history stay inside infrastructure we control.

Requirement: License

What we checked
MIT or similarly permissive, with an active maintainer community.
Why it mattered
We can modify, harden, and redistribute it without renegotiating anything.

Requirement: Model-agnostic

What we checked
Talks to any model through a standard API, including models served from our own cloud account.
Why it mattered
Model choice stays a policy decision, not a product limitation.

Requirement: Pluggable identity

What we checked
Supports SSO (OIDC or SAML) and passes the user's identity through to tools.
Why it mattered
Every action ties back to a real person with real permissions.

Requirement: Tool and connector standard

What we checked
Uses an open tool protocol rather than a proprietary plugin format.
Why it mattered
Connectors we build survive a future runtime swap.

Requirement: Replaceability

What we checked
Clean boundary between the runtime and our configuration, data, and connectors.
Why it mattered
When something better ships next quarter, we can move to it in days.

The Architecture

The deployment has five layers. Only one of them is the open-source assistant.

1. The runtime. The open-source assistant app and its agent loop, deployed as containers in our cloud account, behind our network boundary. No conversation data leaves the account except calls to approved model endpoints.

2. Identity and access. Single sign-on through our identity provider. Roles map to groups we already maintain, so a new hire gets the assistant on day one and loses it the moment they are offboarded. Nobody has a separate password or a personal API key.

3. Connectors with delegated permissions. The assistant reaches company systems such as email, calendar, documents, CRM, and project tools through connectors that act as the signed-in user. If you cannot open a folder yourself, your assistant cannot open it for you. Write actions like sending email or updating records require explicit confirmation.

4. Model gateway. Every model call goes through one gateway we operate. It decides which models are approved, routes requests by task, enforces per-user and per-team budgets, and records token usage and cost. We use enterprise model agreements that exclude our prompts and outputs from model training, and the gateway blocks any endpoint without those terms.

5. Observability and audit. Every conversation, tool call, and model request emits structured logs into our own logging stack. Leadership gets a dashboard with adoption, cost by team, connector usage, and blocked actions. Retention follows our existing data policy, not a vendor’s.

The design principle

Everything that touches company data or company money sits in a layer Metacto owns. The runtime in the middle is swappable.

The Controls That Made Leadership Comfortable

Our leadership team did not ask whether the assistant was smart. They asked four questions. The deployment answers each of them.

What does it cost? Spend is metered per user and per team at the gateway. Budgets alert before they are hit and hard-stop when they are. Expensive models are reserved for the tasks that need them, and routine work routes to cheaper ones.

What can it see? Exactly what the user can see, through the connectors we enabled, and nothing else. Sensitive systems are off by default and turned on by role.

Does our data train someone’s model? No. Approved model endpoints are contractually excluded from training, conversation history lives in our own database, and the gateway rejects calls to any endpoint that is not on the approved list.

Who did what? Every action is logged with the user, the tool, the data touched, and the model used. If something goes wrong, we can reconstruct it.

What We Learned in the First Week

Connectors create the value. A general-purpose assistant with no company context is a nicer search box. The same assistant with access to your calendar, documents, and CRM starts to prepare meetings, draft follow-ups, and pull together account history. The first few connectors matter more than any model upgrade.

Leaders adopt fastest when the assistant knows the business. The most immediate pull came from leadership: meeting prep, board and client briefings, research, and pulling numbers together across systems. That is where context and access pay off most.

Spend visibility changes behavior. Once people could see cost per task, model choice stopped being a debate. Teams picked the cheapest model that did the job well.

Replaceability has to be designed in on day one. Connectors on an open protocol, configuration in our own repository, history in our own database, and models behind our own gateway. That boundary is what lets us move to a better runtime next quarter without starting over.

What Comes Next

We are turning this deployment into a repeatable reference architecture: infrastructure, security boundaries, SSO and role-based access, model routing, connectors, observability, backup and recovery, update process, and support model.

That reference architecture is now a Metacto service. We deploy the same private assistant inside a client’s cloud, connect it to their systems with their approved models, and operate it for them. We are starting with leadership teams, because that is where we saw the clearest value internally.

If you want the reasoning behind that service and the controls it includes, read Managed AI Assistant: Consumer AI Superpowers, With Enterprise Control.

Subscribe to our newsletter

Be the first to get insights on Operational AI, engineering quality, and building systems that move real business metrics.

By subscribing you agree to ourPrivacy Policy.