A model the size of GPT-4 just ran on a $1,600 graphics card. At 100 tokens per second.
Let that land for a second. Qwen 3.8 Flash Next — 125 billion parameters — running on a single consumer GPU at 100 tokens per second. That is not a research demo collecting dust in a lab. That is a setup a serious operator can actually build, rack under a desk, and plug into real business workflows starting this month.
If you have been paying OpenAI or Anthropic by the token and quietly wondering whether there is a smarter way to run this, the answer just got a lot more concrete.
Three Hidden Costs You Are Probably Not Tracking
Every AI workflow you run through a cloud provider carries three costs that rarely show up on the same line of your P&L:
- The API bill. It compounds fast. A support agent handling 10,000 conversations a month, a document pipeline processing contracts all day, an internal ops assistant answering staff questions — those tokens add up to real money, every single month, forever.
- The latency. Every request travels to a data center, gets processed in a queue with everyone else's traffic, and travels back. For real-time use cases — live chat, voice agents, instant document review — that round-trip delay is a user experience problem you are paying for and still not solving.
- The data exposure risk. Your customer records, your supplier contracts, your internal financials — every prompt you send to a third-party API is data leaving your building. Most vendors have solid policies. But "solid policies" and "zero exposure" are not the same thing, and your customers and your legal team know the difference.
Local models eliminate all three. The bill is the hardware, paid once. The latency is a few milliseconds across your own network. And the data never goes anywhere.
What 100 Tokens Per Second Actually Gets You
Speed matters because it determines which use cases are actually viable. Here is a quick calibration: a typical human reads at roughly 200 to 250 words per minute, which is about 3 to 4 words per second. At 100 tokens per second — where one token is roughly three quarters of a word — this model is generating text faster than any human can read it in real time.
That threshold unlocks a specific class of applications that slow local models simply cannot power well:
- Live customer support agents that respond in under a second, maintain full conversation context, and never expose a chat transcript to a third party.
- Document analysts that ingest a 40-page supplier contract, extract the key clauses, flag the risks, and hand you a structured summary before your next meeting.
- Internal ops assistants that your warehouse team, your customer service reps, or your buying department can query in plain English — against your own inventory data, your own order history, your own SOPs.
- Batch processing pipelines that run overnight against thousands of SKUs, tickets, or records without a per-call cost attached to every row.
I run automations exactly like this inside my own import and e-commerce operation. The difference in cost and control compared to routing everything through a cloud API is not marginal — it is structural. Once the hardware is paid for, the marginal cost of every additional query is essentially zero. That changes how aggressively you are willing to automate.
What You Actually Need to Build This
The hardware conversation is simpler than most people expect. A single NVIDIA RTX 4090 — available new for around $1,600 to $2,000 depending on the vendor and configuration — is enough to run a 125-billion-parameter model at production-viable speeds. That is the same class of model that powered early GPT-4 deployments, now hostable on a machine that fits under your desk.
The software layer is open source. Tools like Ollama and LM Studio handle model loading and local inference without requiring you to write a single line of code. Frameworks like n8n let you connect that local model to your existing tools — your CRM, your helpdesk, your spreadsheets, your internal databases — through visual workflow automation that a non-developer can actually manage.
The part that trips most operators up is not the hardware and it is not the software. It is the workflow design: knowing which processes in your specific operation are worth automating first, how to structure the prompts so the model behaves consistently, and how to build the integrations so data flows in and results flow out without manual intervention.
You do not need a data center. You do not need a dev team. You need the right setup and someone who has already built it in a real business.
The Window Is Now
Six months ago, running a 125-billion-parameter model locally was a GPU cluster conversation. Today it is an RTX 4090 conversation. The models are getting more efficient faster than the hardware is getting more expensive. The operators who figure out local AI infrastructure now — who build the workflows, train their teams, and stop paying per token — will have a structural cost and speed advantage over the ones who wait for this to feel more "mainstream."
It already is mainstream. It just does not look like it yet from the outside.
If this opened your eyes to what is actually possible right now, the next step is straightforward. At Maqia, we help small and mid-sized businesses design and deploy local AI workflows — the kind that plug into your actual operation, not a generic demo. No code required on your end. If you want to understand exactly how a setup like this would work for your business, visit maqia.co and book a call. We will look at your workflows, identify where local AI makes the most sense, and show you what a real build looks like.