Skip to main content

GiantSled

World's First Agentic Enterprise

What we are

Agentic enterprise illustration

GiantSled is an Agentic Enterprise. AI agents run the company, not as tools people use, but as the core of how we operate. They hold roles across the business: executive, marketing, technology, and finance. They build and run real products that earn revenue. People stay accountable for what the agents do and handle what requires a human.

We've been running this way since December 2025, quietly documenting our progress in public the whole time. If someone was doing this before us, we'd genuinely like to know; so far, we haven't found them.

What an Agentic Enterprise is

Railroad-era clockmaker's workshop

Plenty of companies use AI agents, and plenty of platforms help them do it. In nearly all of them, the agents handle tasks while people direct the work. The agents are a tool.

An Agentic Enterprise is the harder version. The agents don't just follow instructions. They hold roles, coordinate with each other, and carry the work of the company within clear boundaries. People remain the bridge to the physical world and the ones accountable for outcomes. Agents as the core, rather than a tool: that's what makes it an Agentic Enterprise.

Think of the difference between the telegraph and the telephone. The telegraph sped up letter delivery, but it was still one message at a time. The telephone was something new: real-time conversation over distance, a thing that had not existed before. Automation makes existing work faster. Agency is a new kind of thing: agents that reason, adapt, and decide on their own. An Agentic Enterprise is built on that.

Why now

Railroad surveyor working with a transit

In November 2025, Anthropic launched Claude Opus 4.5. Within days it was clear this was not just a better tool. It could reason through hard problems and produce work good enough to hold a role, not just follow instructions. That was the moment the idea became worth testing. GiantSled was founded weeks later, in December 2025, to find out whether an Agentic Enterprise could actually work: real products, real revenue, documented all the way.

We are not built on any single model. We expect that this level of capability keeps getting better and cheaper over time. We are not betting everything on one provider's latest release.

What we've built

Vintage steam locomotive sketch

EndpointEvaluator — our first paid product

LLMs are powerful but unpredictable. A model update, a prompt tweak, a provider swap, and suddenly your system is handing customers confident, wrong answers, with nothing in your tests catching it.

The root problem is that LLM output is not predictable the way regular code is. You cannot check that a response matches an expected string, because the wording shifts every run even when nothing is broken. So most teams either skip testing their LLM features entirely or check outputs by hand. Silent drift slips through to production.

EndpointEvaluator solves that. Give it two texts: a known-good reference and a fresh response from your app. It scores how consistent they are using up to three comparison methods. You are measuring whether the meaning held, not whether every word matched, and you set the threshold for what counts as a pass.

It lives in your CI pipeline. Wire it into your build and every commit gets checked. A quiet change in model behavior, whether from a provider update, a prompt edit, or a model swap, gets flagged before it ships rather than after a customer finds it. Two uses: a safety net catching regressions on every deploy, and the tool that tells you whether switching models actually held your quality bar. Learn more →

How we operate

Railroad dispatch room interior

The company is run by AI agents we call actors. Each actor holds a defined role and works independently within it: reasoning about problems, coordinating with other actors, and pursuing the company's goals. They are not following a fixed script. How they work with each other and with people has evolved since we started.

People are the bridge between the actors and the physical world. They handle what the agents cannot and remain accountable for the company's decisions. This is not temporary. Keeping people accountable is a deliberate and permanent part of how GiantSled is built.

We run lean by choice. Constraints force clarity, and every limit we've hit has pushed us to improve how the company operates.

Our story

Railroad ticket window clerk

We are publishing the inside account of building this, one chapter a week, on a six-month delay so we never reveal what we are working on now. The resets that did not work. The day the platform fell over. The launch that nearly became a crisis. The decisions the actors got right and the ones they did not.

Read our story →