Anant Shrivastava

Anant Shrivastava

@anantshri

Owning Your AI: Building a Private Stack You Can Actually Live With

Submitted Oct 1, 2026

One-line summary

This session documents the journey of buiding and operating AI / LLM systems on hardware you directly control be it home or office. The key focus is on model choices, trust, power, isolation, maintainability and real-world usefulness.

What problem are you addressing?

A lot of private AI discussions either focus on using hosted APIs or jump straight to datacenter-scale infrastructure.

This talk is about the space in between: running AI on hardware you directly own and control.

That might be a single workstation with a GPU, an Apple system with unified memory, a small collection of machines connected over Tailscale, or an isolated box running an agent that should never send its context outside your network.

The constraints are very different at this scale.

You care about things like whether the machine remains usable while inference is running, whether a model can stay loaded all day, how much heat it produces, what happens to power consumption, whether a smaller model is intelligent enough for the task, and whether an upgrade is worth the disruption.

AI/LLM APIs are getting cheaper per token, while the number of tokens consumed per task is often increasing. People sometimes move toward self-owned AI primarily as a cost-saving measure. I would argue against treating cost as the main reason.

Self-owned AI is often more expensive, slower, and harder to maintain than calling a good API.

The value is control: control over models, data, network boundaries, runtimes, tools, upgrades, agents, and eventually the complete stack.

Who is the intended audience?

Platform Engineering, SRE, DevOps, Infrastructure, Developers, Security Engineers, AI/ML practitioners, and technically curious people experimenting with self-owned AI on hardware they directly control.

The focus is not datacenter-scale AI infrastructure. It is the smaller end of the spectrum: a workstation, home lab, research setup, small office environment, or a few systems connected together to provide private inference and agent capabilities.

Level

Intermediate.

Some parts will go deeper into inference runtimes, quantization, model behaviour, hardware optimization, networking, supply-chain risks, and agent isolation.

The talk assumes that attendees are comfortable with infrastructure concepts, but they do not need prior experience operating local LLMs.

Practical takeaways

  1. A practical framework for selecting models based on how useful they are on the hardware you actually own, considering intelligence, heat, power, memory pressure, responsiveness, tool use, and maintainability rather than simply model size or benchmark numbers.

  2. An approach for building self-owned AI where models, runtimes, networking, routing, prompts, trust, airgapping, tools, and agent workflows are treated as parts of an infrastructure stack that must be tested, constrained, pinned, and maintained.

What will you share?

I will share architecture, experiments, failure modes, open-source tooling, and lessons from running private AI across different hardware and software stacks at workstation, home-lab, and small-office scale.

The talk will cover:

  • Why cost alone is usually a poor reason to move from hosted APIs to self-owned AI
  • What control and ownership provide in return for accepting additional cost and operational complexity
  • Why “can this model run?” is different from “should this model run here?”
  • Choosing models based on workload rather than leaderboard position
  • Why intelligence does not map cleanly to model size
  • Choosing between model sizes, quantization levels, dense and MoE architectures
  • Power, heat, responsiveness, fan noise, and system degradation as real operational constraints
  • Why highest tokens/sec or prefill performance is not always the right optimization target
  • Experiments optimizing llama.cpp for inexpensive Intel Arc Pro B70 hardware, and what that taught me about useful performance versus benchmark performance
  • Prompt and context optimization experiments with Hermes and smaller models before assuming that a larger model is required
  • Whether newer models are actually better for the task you care about
  • Building your own regression tests and golden tasks for model upgrades
  • Establishing trust through model provenance, version pinning, behavioural testing, and controlled tool use
  • Running airgapped or highly restricted agents
  • Using Tailscale and similar approaches to connect selected systems while keeping other services isolated
  • Deciding what an agent should be able to see, what it should connect to, and what should remain outside its trust boundary
  • Deciding what belongs in the model and what should come from retrieval, tools, or APIs
  • Fine-tuning or adapting smaller quantized models for narrow use cases
  • Using, customizing and owning llama.cpp and llama-swap as parts of the inference stack
  • Modifying open-source runtimes when the default behaviour does not match the hardware or workload. Deciding if its specific to you or needs to be upstreamed.
  • Experiences comparing coding-agent approaches such as Claude Code and OpenCode, including where model quality helped and where the surrounding harness mattered more
  • Using AIDC as an example of moving control into the harness, including context management, references, execution, review, and lifecycle management
  • Supply-chain concerns around models, runtimes, Python packages, containers, plugins, agent skills, and other dependencies entering a supposedly private environment
  • Where local models fail, where APIs remain better, and where a hybrid architecture is the more sensible answer

The individual projects and tools are examples. The aim is to extract operational patterns that remain useful even when models and frameworks change.

What is your experience with this problem?

  • Internal engineering project
  • Open-source project
  • Experiment or prototype
  • Research / investigation
  • Hard-earned engineering lesson

Over the last few years I have been running public and private AI models side by side across NVIDIA, Intel, and Apple hardware.

Most of this work happens at the scale this talk is aimed at: individual workstations, home-lab systems, research machines, and small office infrastructure rather than datacenter deployments.

My work has included:

  • Maintaining a customized llama-swap setup for my own model routing and lifecycle requirements
  • Contributing performance and hardware-specific changes to llama.cpp, including optimization work around Intel Arc Pro B70
  • Experimenting with model sizes, quantization, speculative decoding, context sizes, prompt optimization, tool calling, and different inference backends
  • Comparing larger and smaller models on the same workloads rather than assuming parameter count determines usefulness
  • Experimenting with Hermes Agent and smaller models to understand how much capability can come from better prompts, tools, and context
  • Comparing coding-agent environments such as Claude Code and OpenCode to understand where capability comes from the model and where it comes from the harness
  • Running local, isolated, and selectively connected agents
  • Using Tailscale to connect selected AI systems without exposing the complete environment
  • Building private AI infrastructure for coding, security research, and automation
  • Maintaining security and engineering tools where AI is used as part of development, analysis, and review
  • Building AIDC as an agent-oriented development environment with explicit context, references, execution, and review workflows
  • Operating both local models and hosted APIs depending on the workload

This has been less about building one polished “AI platform” and more about continuously changing the stack until it behaves the way I need and owning the resultant stack.

What approaches failed, disappointed, or created unexpected problems?

Several things turned out to be less useful than they first appeared.

  • Chasing raw tokens/sec often improved benchmark numbers without improving the actual experience of using the system.
  • A model can technically fit in memory and still be the wrong model for the machine. Large models sometimes made the rest of the system unpleasant to use because of heat, memory pressure, loading time, fan noise, or sustained power draw.
  • Model size also turned out to be an unreliable shorthand for intelligence. For some workloads, a smaller model with better prompts, cleaner context, and better tools was more useful than a larger general model.
  • Newer models were not always better. Some improved reasoning while regressing in tool calling, instruction following, structured output, or behaviour under quantization.
  • Quantizing more aggressively sometimes saved enough memory to make a model fit while damaging exactly the capability I cared about.
  • Long context introduced memory and KV-cache costs that were easy to ignore while looking only at model size.
  • Coding-agent experiments also showed that changing the model does not automatically fix the workflow. Comparing systems such as Claude Code and OpenCode made it increasingly clear that context management, execution controls, tool design, review loops, and the surrounding harness can matter as much as the underlying model.
  • Airgapping turned out to be more complicated than disconnecting a machine from the internet. Models, runtimes, packages, documentation, tools, updates, and agent capabilities all need controlled ways to enter the environment.
  • This creates another problem: supply chain. A system can be completely private while running and still depend on externally produced models, packages, containers, plugins, agent skills, and binaries.
  • Even deciding what to expose became an architecture problem. Giving an agent access to everything is easy but creates a very large trust boundary. Giving it access to nothing produces a very private but mostly useless system.
    The biggest lesson has been that self-owned AI is much more of an infrastructure problem than an application problem.

What will you do differently today?

  • I now optimize around the workload rather than the model.
  • I prefer a model that is good enough, predictable, and comfortable on the machine instead of constantly replacing it with whatever is newest.
  • Before moving to a larger model, I look at whether the problem can be solved by improving prompts, reducing unnecessary context, providing better tools, or improving the agent harness.
  • I test model upgrades against my own workloads before adopting them.
  • I separate model intelligence from tool-provided knowledge. If accurate information can be retrieved from a deterministic source or API, I prefer teaching the model when and how to use that source instead of trying to embed all knowledge into the model.
  • I increasingly treat the harness as its own engineering problem. A model is only one component in an agent system.
  • I treat model files, prompt templates, inference runtimes, routing configuration, tool definitions, and agent behaviour as versioned parts of the stack.
  • For networking, I explicitly decide which systems should communicate. Tools such as Tailscale make it possible to connect selected nodes without exposing inference endpoints or agents to the public internet.
  • For agents, I prefer restricted environments with explicit tools, limited network access, and clear separation between reasoning and execution.

What trade-offs did you consider?

The largest trade-off is local versus hosted.

  • Hosted APIs are often faster, easier to operate, and provide access to more capable frontier models.
  • Self-owned AI provides control over data, model versions, routing, network boundaries, tools, upgrades, and lifecycle.
  • I do not treat this as an either-or decision. I use both. The goal is to decide what belongs where.

Other trade-offs include:

  • Larger model versus smaller specialized model
  • Model size versus actual task intelligence
  • Better model versus better prompt and context
  • Model capability versus harness capability
  • Q4 versus Q8 versus higher precision
  • Always-loaded model versus load-on-demand
  • Dense versus MoE
  • Raw throughput versus heat and power
  • Peak benchmark performance versus stable daily operation
  • Private model versus public API fallback
  • Strict airgap versus controlled egress
  • Broad agent access versus narrowly scoped tools
  • Convenience of connectivity versus size of the trust boundary
  • Building custom tooling versus adapting existing open-source systems
  • Pulling the latest model or runtime versus keeping a known-good version
  • Convenience of third-party components versus the supply-chain risk they introduce

The main thing I optimize for is usefulness over time on hardware I actually own, not peak performance during a benchmark.

How can this help other practitioners?

The session should help practitioners avoid unnecessary hardware purchases, benchmark chasing, and architectures where every problem is solved by loading a larger model.

The main patterns I want attendees to take away are:

  • Treat self-owned AI as infrastructure
  • Choose models based on the job they need to do
  • Measure behaviour on your own hardware and workloads
  • Do not use parameter count as a substitute for testing intelligence
  • Improve prompts, context, tools, and harness design before assuming that the answer is a larger model
  • Test model upgrades like software upgrades
  • Pin what works
  • Understand your power and thermal limits
  • Use smaller models when they are good enough
  • Train or adapt behaviour for narrow use cases instead of trying to put all knowledge into the model
  • Use tools for facts and deterministic work
  • Treat the agent harness as part of the platform
  • Explicitly decide what agents can see, access, and communicate with
  • Treat models, runtimes, containers, plugins, and agent tools as part of your software supply chain
  • Keep agents constrained
  • Use hosted APIs where they genuinely make more sense
  • Own only as much of the stack as you are willing to maintain

The goal is to help people build small, self-owned AI systems that are useful enough and boring enough to depend on.

Current state

In progress, with open-source components and daily use across personal, research, home-lab, workstation, and small-office infrastructure. The linked projects are active works in progress and change regularly because I use them as daily drivers.

Also, I did a run of a tangential topic for security focused audience. That talk was focused on giving overview and moving towards security of AI. he proposed talk is a much deeper insight into what I have been building, operating, and measuring, including the results and trade-offs.

Tags

#privateai #selfhostedai #inference #agents #homelab #mlops #platformengineering #devops #sre #security #llama #quantization #airgapped #opensource #power #hardware #supplychain #tailscale #codingagents #casestudy #workinprogress

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy