In the past couple of months, we’ve gotten the opportunity to talk about self-hosted AI with many different teams and organizations. Yesterday, we had an awesome event about self-hosted AI where we got even more of those questions (a recording will be shared soon!).
Across this, two specific questions (and their many sub-questions) kept popping up again and again - so we decided to gather them, and answer them here.
- How do I get started working with self-hostable AI?
- What does self-hosted AI cost (in actual numbers)?
- Why we bought a GPU anyways
- Conclusion
How do I get started working with self-hostable AI?
A quick way to get started tinkering with local models (especially for chat) is to download a free product which ships with inference and model-downloading built in, such as:
- LM Studio (free, but not open-source)
- OpenWebUI (open-source, but a little more technical in its setup)
If you have access to a GPU or a decently chunky MacBook, you can run the models right there on your machine and have a fully local setup. See Picking the right model below for more info on model selection.
Product integration of AI
If you’re more inclined towards integrating an AI model into a software stack, your best bet is to download an inference engine (the “API wrapper” of the model) directly, downloading the model through that and exposing it as an API on your machine.
There are two main inference engines we would recommend:
- llama.cpp for personal use
- vLLM for concurrent use by multiple users
llama.cpp is super quick to get up and running, while vLLM gives you a lot more of a production-grade experience.
Picking the right model
It’s hard to conclusively give a perfect answer to what model you should run, as it depends on use-case and hardware. But for some general guidelines:
- If you have access to around 32GB of RAM, Qwen 3.6:35B-A3B is a good candidate
- If you have access to 24GB VRAM or more, Qwen 3.8:27B is a good candidate, and is slightly more intelligent and better at tools than the above. This was the model you saw yesterday.
- If you have access to 96GB VRAM or more, you have options like Qwen3.8-Flash-Next or DeepSeek-V4.1-Flash
This is where you can really start competing with the big frontier models - although we have yet to meet a single customer who found that core model intelligence was their actual issue past the above models. Context and integration is king!
And in general; it’s better to have a smaller model for a specific problem, than a generic model for all of them (especially if it handles problems you don’t even have).
What does self-hosted AI cost (in actual numbers)?
There are a few important points to consider when it comes the business case of self-hosted AI. It’s important to separate the product, the data, the integrations etc. from the fact that at the end of the day, you need to pay for compute somewhere. This is also true, if you’re currently on a Claude/ChatGPT subscription - you’re just paying for a bundle of both product and compute.
The actual software product that you use to interact with AI is usually cheap and simple to self-host, like any other software you might have been used to. This leaves the expensive part: the compute.
When it comes to self-hosted AI, we have a lot more options for buying compute. Concretely, the three most obvious options are:
- Buying hardware directly
- Pro: full ownership, one-time cost (return of investment more clear)
- Con: big upfront cost
- Renting compute time directly (e.g Hetzner, RunPod etc.)
- Pro: lower upfront cost, easy to test various hardware options directly (e.g before buying)
- Con: continuous costs (and rising ones at that), external provider for inference (less privacy)
- Subscription with a cloud provider (e.g Azure, AWS, GCP etc.)
- Pro: lower upfront cost, might be someone you already share data with, easy to scale (at the click of a button and the swipe of a credit card)
- Con: continuous costs (and rising ones at that), external provider for the entire stack (a lot less privacy and control)
To give you a sense of some concrete numbers, let’s imagine we are trying to run a Qwen 3.8:27B model. This is a pretty powerful model, which requires around 32GB of VRAM to use comfortably - and a model we frequently use with customers.
We’re going to target an RTX5090 GPU, which:
- Fits the model
- Is extremely fast (faster than ChatGPT)
- Can support a throughput allowing a team of a 2-4 people (conservative estimate) using it concurrently
- This is a bit of a generalization, it depends a lot on their activity and the type of work they do. If you’re just asking questions and reading documents on and off during your day, you probably support double that. If a concurrency bottleneck is hit, requests are simply queued.
Example: The business case of a Qwen 3.8 model on an RTX5090
Let’s assume we have the above team of 4 people, and that they’re currently using Claude’s Opus 5.5. At the time of writing, this model costs $20 (€17.56) per 1M tokens output.
An RTX5090 can output 100 tokens/sec or 360.000 tokens/hour (for a single stream). Multiple concurrent users can absolutely share a GPU, but this is much more governed by their usage patterns, so we will stick to single-user calculations for now.
To generate 1M tokens, our RTX5090 would have to run for 1.000.000 / 360.000 = 2,78hrs.
Let’s use this as a base to compare our three options for self-hosted AI.
Buying an RTX5090
An RTX5090 currently costs around €5.000. The upfront cost is most an expense/return-on-investment question. Let’s start with comparing on-going token pricing.
Running your own GPU only has one cost; power. The RTX5090 requires 600 watts at peak load. In Denmark, a conservative estimate of 1 kWh is roughly DKK3 (€0.4).
This makes it fairly easy to calculate the cost of 1M output tokens:
2.7hrs * 600/1000 (watts/hr) * €0.4 (price-per-kWh) = €0.648
This means you would save €17.56 - €0.648 = €16.9 per 1M tokens generated.
However, there are of course two obvious caveats:
-
You’d be buying a €5.000 GPU. The ongoing savings would offset this after
€5000 / €16.9 = 296M output tokens. This is roughly equivalent to 34 days of non-stop peak use of the RTX5090. -
Qwen 3.8 and Opus 5.5 are different models, and there is no doubt that Opus is more intelligent. While we have yet to see Qwen 3.8 not be intelligent enough for a job at hand, one could argue that that is what you’re paying for. For an alternate comparison, Claude Haiku 4.5 (their cheapest model) costs €4.39 per 1M output tokens, and would have a RoI of ~
1.349B tokens (151 days of non-stop peak use).
Note that in real-world use for an actual team, both of the following are true:
- The GPU is rarely running at peak usage non-stop. This means you pay less for power than the above assumes, but also that the RoI could be longer in reality.
- We hinted at a team of 2-4 people using it above (and multiple streams). Each user would not require their own GPU, but we didn’t want to involve this in the above calulations. If your workload allows 2 people running concurrently on the RTX5090 (which in our experience is very feasible), the business case looks twice as good.
Renting compute on an RTX5090 on RunPod
Renting the same setup from RunPod is a lot simpler to calculate - at the time of writing, renting an RTX5090 for 1hr costs €0.88.
This means that generating 1M tokens (which takes 2.7hrs) costs €0.88 * 2.7 = €2.38.
Running Qwen 3.8:27B with a cloud provider
It is borderline impossible to get clear prices for open-weight models in Azure Foundry, Amazon Bedrock and GCP. Their prices vary by model type and region.
The only publicly available price we could find was from Qwen itself, which takes €2.63 per 1M output tokens. This puts it just above renting compute, which doesn’t seem too far fetched.
Business case comparison summary
Summarizing the above three cases, we get the following results:
| Case | Cost pr 1M tokens | Upfront cost |
|---|---|---|
| base: Cloud (Opus 5.5) | €17.56 | None |
| 1. Buying | €0.648 | € 5.000 |
| 2. Renting (RunPod) | €2.38 | None |
| 3. Cloud (Qwen) | €2.63 | None |
So at a glance, it looks cheap to buy the GPU. But the upfront GPU cost is of course the downside - and it’s a cost you’ve paid regardless of usage.
Comparing buying hardware to the other three options, you’d need to generate the following amount of tokens to reach return-on-investment:
- 2886M tokens across 324 days of peak usage compared to renting on RunPod
- 2523M tokens across 292 days vs. the Qwen cloud price
- 296M tokens across 34 days vs. the Opus 5.5 Anthropic price
On the other hand, when buying a GPU you know exactly what your tokens will cost and that price will remain static over the lifetime of the hardware (generally estimated to be around 5 years by professionals)
This aligns with our general recommendation: if cost is your primary measure (ignoring privacy and control), it is most likely cheaper to rent hardware or to buy tokens from a cloud provider, than it is to buy your own hardware (at least for now). It doesn’t help that the market is so volatile that GPU prices are doubling every 6 months, whilst cloud providers are slower to increase their prices (although they are increasing).
NB: It’s important to note that many of the values of using open-weight AI (behaviour control, system prompt, tools, deciding when it changes etc.) stays consistent across the three business cases, aside from privacy, which of course degrades from option 1 to 3. The rest is purely an exercise in cost-comparison.
Why we bought a GPU anyways
So if it’s more expensive, why did we decide to buy ours anyways? First off, it was important for us to have the flexibility to host the models we wanted. As you can see, we couldn’t even find this particular model anywhere with any of the big cloud providers. It’s just not available. It also makes it easy for us to test new models, since we can just spin them up on the hardware. We don’t need to wait for the providers to do so.
Second of all, we wanted to provide our customers a consistent pricing. Usage-based pricing might be cheap right now, but we have already seen prices increase as the era of subsidized AI is moving towards its end. Furthermore, with everything happening around the world, prices increasing, and the many different ways we could imagine using AI (now and in 6 months, a year etc.), this gives us a clear-cut platform to base our pricing on. We can offer customers a single monthly cost, knowing that our own costs won’t balloon because a third party increased theirs.
One point to add here: we’ve come to appreciate that we don’t have to think of pricing all the time, now that we have the hardware. We have a transcription model, so now we want to just transcribe all of our awesome conversations going on at the office. That’s something that brings value to us, because we can easily search among that. But if we had to think: “hmm, how much does it cost to transcribe useless conversations?” we probably wouldn’t even consider it. But with our own hardware, experimenting is practically free. It has already led to some awesome internal tooling that now helps us on the daily.
Lastly, we’re back to privacy. We can’t even begin to consider transcribing our daily conversations at the office, if we had to give all of that data to a third party. Sometimes we talk about customers and their proprietary information, that we’re not allowed to share. We also like recording customer conversations for future analysis later and the answer to the question: “who gets this data?” is now always “you’re looking at them”. The same goes for text that we generate with our model: we couldn’t talk with our internal agents about proprietary customer cases, if we used external third parties to process the information. Our customers have never accepted us doing that and doing it anyways, would be completely irresponsible.
Conclusion
There isn’t necessarily a right or wrong answer in all of the above. However, this post has hopefully managed to demonstrate that with self-hosted AI you get a choice. If you’re using one of the AI providers, you’re already on one of the above business models, whether it’s clear to you or not - the choice just lies with the AI provider you’ve gone with.
For us, one of the goals of working with self-hosted AI is to separate the decisions of product, data, integration etc. from the simple business case of where you compute runs at the end of the day. That choice should be purely financial and should be easy to change.
So while we cannot make the call for you, we hope this article at least gave you some insights into the options you have with self-hosted AI. And that we at least helped you ask and answer one very clear question: What is the price of owning my AI stack and its data?