The Shelled Thesis
Intelligence Is a System
The full stack for useful AI.
OverviewAI becomes useful through the system around it
AI becomes useful through the system around it: where it runs, what it knows, how it works, and who stays in control. Shelled is building the shell that brings that system together. Our direction spans local and cloud compute, model serving and adaptation, shared context, collaborating agents, and the applications people use to get work done.
Our thesis is simple: the next leap in useful AI will come from improving the whole stack. Better models expand what is possible. Better systems make that capability affordable, reliable, and available where people need it.
More useful work from every token, dollar, watt, and minute.
Run intelligence where it makes sense.
Power, hardware, networking, inference software, and access to models. Make local, on-premises, and cloud resources work together so placement follows the workload.
Turn available capability into completed work.
Connect planning, research, implementation, verification, and follow-through — giving each step the models, tools, context, and resources it needs, within clear budgets and permissions.
Make each completed task inform the next.
Every workflow produces evidence. With permission to use it, routing, memory, workflows, evaluations, and models adapt — evaluated against held-out tasks and real constraints.
01 — The ShellThe shell around intelligence
A shell gives a system structure, protection, and an interface to the world. For Shelled, that means connecting intelligence to the resources, context, tools, and permissions it needs to act.
A task has a goal, a budget, a deadline, and a standard of quality. Meeting those requirements involves decisions across the stack: which model to use, where to run it, what information to provide, how to divide the work, when to check a result, and when to involve a person.
We believe those decisions belong in a coordinated system. Each should reflect the needs of the task and the constraints of the person or organization using it. That is the role we see for Shelled: a common environment for directing work across models, machines, and agents, with visibility into what happens and control over what happens next.
02 — Three PillarsThree pillars hold up the stack
The Shelled thesis rests on three pillars — Infrastructure, Orchestration, and Improvement. Each addresses a different question: where intelligence runs, how it completes work, and how it gets better over time.
Pillar One — InfrastructureRun intelligence where it makes sense
AI needs a practical foundation: power, hardware, networking, inference software, and access to models.
Choices at this level shape the cost, speed, privacy, and availability of everything above it. Our approach is to make local, on-premises, and cloud resources work together. A team should be able to keep sensitive workloads on its own machines, use available capacity for recurring work, and draw on cloud services when a task needs additional capability or scale.
Placement should follow the workload. The decision must account for model quality, memory, latency, utilization, data requirements, and total operating cost. Local capacity has hardware and maintenance costs; cloud capacity has service and usage costs. The useful comparison is the cost of completing the work at the required quality.
The longer-term opportunity extends into how compute is provisioned and powered. We want intelligence to be practical at the scale of a personal device, an office, and a data center, with infrastructure choices that serve the work running above them.
Pillar Two — OrchestrationTurn available capability into completed work
A goal can require planning, research, implementation, verification, and follow-through. Orchestration connects those steps.
We are building toward persistent workflows in which agents can collaborate, preserve progress, check results, recover from failures, and request human judgment when needed. Model selection and infrastructure placement are part of that same process.
For a software task, this could mean a local model handling routine edits, another model reviewing a difficult design decision, and tools testing the implementation. Shared state keeps those activities connected to the original requirements. Permissions and budgets define the boundaries of execution.
Coordination must earn its overhead. Adding agents can add latency, duplicate work, and increase token consumption. A useful orchestrator should recognize when a single model call is sufficient, when parallel work helps, and when more reasoning or review is worth its cost. Our objective is dependable completion with an appropriate amount of computation and human attention.
Pillar Three — ImprovementMake each completed task inform the next
Every workflow produces evidence: what succeeded, what failed, which choices were expensive, and where people had to intervene.
With permission to retain and use that evidence, a system can become better suited to the work it performs. Improvement can happen at several levels. Routing can select a better model. Memory can preserve a useful decision. A workflow can remove a redundant step. Evaluations can catch a recurring mistake. Where justified by the workload and available data, model adaptation, fine-tuning, or training can build more specialized capability.
These changes should be evaluated against held-out tasks and real operating constraints. We want evidence that a change improves completion quality, cost, speed, or reliability before it becomes the default. Versioning and rollback make that process inspectable and reversible. This is a practical path toward systems that improve through use, with people setting the objectives and deciding which changes to accept.
03 — The Scope of the StackThe pillars connect through a shared architecture
| Layer | Role in the Shelled thesis |
|---|---|
| Energy and physical infrastructure | Understand and improve the power, cooling, connectivity, and deployment conditions that support AI. |
| Compute and inference | Coordinate local devices, owned servers, and cloud capacity around workload needs. |
| Models and adaptation | Select, serve, evaluate, and adapt models for the work they perform. |
| Context and memory | Keep relevant knowledge and task state available within clear access and retention boundaries. |
| Agents and orchestration | Plan work, use tools, coordinate execution, verify results, and manage recovery. |
| Applications and interfaces | Give people practical ways to direct AI and use its output. |
Building across this stack means understanding the interactions between its layers and integrating strong components through open interfaces. Teams should retain the freedom to change models, providers, and deployment environments as their needs evolve.
04 — TrustTrust across every layer
Trust belongs in the architecture. People need to know what a system can access, what it is doing, what it costs, and which actions require their approval.
Our design principles are open interfaces, observable execution, scoped permissions, explicit budgets, and meaningful human oversight. As workflows become more autonomous, those controls should become more capable and easier to use. The ambition is AI that organizations can inspect, direct, and improve while retaining control of their data, infrastructure, and decisions.
05 — Where We StartStart with software. Build for broader work.
Software engineering is our starting point. It combines complex reasoning with concrete artifacts, tools, and opportunities to test whether an outcome meets its requirements.
The same foundations can support other forms of work: research, analysis, internal operations, and specialized enterprise workflows. Expansion should follow evidence that the system can meet each domain's quality and control requirements. We intend to build that broader foundation through practical products. Each application should help us understand which improvements matter across the stack and which capabilities need deeper investment.
06 — Measure Useful WorkWe judge progress by outcomes
We judge progress by outcomes that meet an agreed standard. The central economic measure is cost per accepted task, including compute, retries, infrastructure, and human review.
We also care about time to an accepted result, completion rate, defects, intervention required, and energy per task where it can be measured. Comparisons should use equivalent workloads and quality requirements. Token savings matter when they improve those outcomes.
07 — The Long-Term HorizonReady for advances in intelligence
Increasingly capable AI, including potential AGI and ASI, belongs within our long-term vision. The architecture should be ready to put advances in intelligence to work while preserving the ability to evaluate, govern, and choose how they are used.
Our immediate work is concrete: connect the stack, improve its economics, and help people accomplish more with the intelligence available to them. Shelled's ambition is to make increasingly powerful AI useful at every scale, from the machine on your desk to the infrastructure behind an organization.