<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Night thoughts]]></title><description><![CDATA[Unfiltered late-night thoughts on software engineering, system architecture, and technical problem-solving. Brain dumps, code logs, and reflections from the quiet hours.]]></description><link>https://nightthoughts.me</link><image><url>https://cdn.hashnode.com/uploads/logos/6a92730f9a9aa7f72e74fdf4/30bf9d05-14bc-4b25-99d8-950183709c80.png</url><title>Night thoughts</title><link>https://nightthoughts.me</link></image><generator>RSS for Node</generator><lastBuildDate>Sun, 04 Oct 2026 22:37:58 GMT</lastBuildDate><atom:link href="https://nightthoughts.me/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[How a Laptop Learned to Run a Giant]]></title><description><![CDATA[New to the jargon? Start with the glossary →

A 70-billion-parameter model is roughly 140 GB of data. A mid-range gaming laptop has between 8 and 16 GB of video memory.
That's like trying to pour a sw]]></description><link>https://nightthoughts.me/how-a-laptop-learned-to-run-a-giant</link><guid isPermaLink="true">https://nightthoughts.me/how-a-laptop-learned-to-run-a-giant</guid><category><![CDATA[AI]]></category><category><![CDATA[hardware]]></category><category><![CDATA[local ai]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[llm]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 04 Oct 2026 10:27:15 GMT</pubDate><content:encoded><![CDATA[<p>New to the jargon? <a href="https://nightthoughts.hashnode.dev/plain-english-glossary-ai-terms-for-the-rest-of-us">Start with the glossary →</a></p>
<hr />
<p>A 70-billion-parameter model is roughly 140 GB of data. A mid-range gaming laptop has between 8 and 16 GB of video memory.</p>
<p>That's like trying to pour a swimming pool into a teacup. It shouldn't fit. So how is it running?</p>
<p>The answer is that we stopped trying to force the whole model into memory at once. Instead, we shrink it, we share the load, and we only use the parts we actually need. Here are the three tricks that make it possible.</p>
<h2>Trick 1: Shrinking the Model</h2>
<p>The first problem is pure size. You can't fit 140 GB into an 8 GB space. So we make the model smaller.</p>
<p><strong>Quantization</strong> is the biggest lever. Think of it like a high-resolution photo. The original model uses 16-bit precision (called FP16), where every single parameter takes up 2 bytes of memory. That's the "high-res" version. Quantization drops that down to 8-bit (INT8), then 4-bit (INT4). At INT4, every parameter takes just 0.5 bytes.</p>
<p>That 140 GB model suddenly drops to around 35 GB. You lose a little bit of quality—the photo gets slightly blurrier—but you save an enormous amount of memory.</p>
<p><strong>Pruning</strong> is the second lever. Researchers look at the model's billions of parameters and find the ones that aren't pulling their weight. They zero them out, effectively deleting them. It's like trimming dead branches off a tree so the rest can breathe.</p>
<p><strong>Knowledge distillation</strong> is the third. Instead of compressing a big model, you train a small model to copy a big one. The small model learns from the big model's answers and reasoning, like an apprentice learning from a master. You end up with a model a fraction of the size, carrying a surprising amount of the original's capability.</p>
<h2>Trick 2: Sharing the Load</h2>
<p>Shrinking helps, but a 35 GB compressed model still doesn't fit comfortably into 16 GB of video memory. So we use everything the computer has.</p>
<p><strong>Unified memory</strong> changed the game on Apple Silicon. In a traditional PC, the CPU has its own RAM, and the GPU has its own VRAM. Moving data between them is slow. Apple's M-series chips use a single pool of memory shared by both the CPU and the GPU. A MacBook Pro with 128 GB of unified memory can hold a ~65B model comfortably—no copying required.</p>
<p><strong>CPU offloading</strong> is the PC equivalent. Instead of demanding that the entire model live on the graphics card, tools like <code>llama.cpp</code> split the model up. The "hot" layers that are needed right now stay on the GPU. The rest sit in system RAM, and the CPU handles them.</p>
<p>Is it slower than pure GPU inference? Yes. But it's a thousand times better than not running at all.</p>
<h2>Trick 3: Doing Less Work</h2>
<p>The final trick is architectural. Modern models are simply smarter about how they use themselves.</p>
<p><strong>Mixture of Experts (MoE)</strong> is the biggest shift. Instead of a monolithic model that uses all 70 billion parameters for every single question, an MoE model is made of many smaller specialist teams. When a question comes in, a router sends it to just a few of those experts—maybe 8 to 10 billion parameters' worth. The model is huge on paper, but computationally cheap in practice. Examples like Mixtral, Qwen2.5-MoE, and Grok use this to deliver massive capability without massive compute costs.</p>
<p><strong>Sparse attention</strong> reduces the memory and compute required for long contexts. Instead of forcing the model to look at every single word it has ever seen, sparse attention focuses only on the relevant parts. It cuts the math down drastically.</p>
<p>And then there's just better design. Models like Llama 3 and Qwen 2.5 are simply more parameter-efficient than their predecessors. A smaller model today beats a bigger model from two years ago.</p>
<h2>The Reality Check</h2>
<p>So what does this actually look like on your desk?</p>
<p>If you have a <strong>Mac Mini (M-series) with 32 GB of RAM</strong>, you can run Qwen2.5-14B or Llama-3-8B comfortably at INT4. You might even squeeze in a 32B model if you're aggressive with quantization and don't have much else open.</p>
<p>If you have a <strong>MacBook Pro with 64 GB or 128 GB of unified memory</strong>, the world opens up. You can run a 70B model at INT4. It's slow—roughly 5-10 tokens per second—but it runs, completely offline, no API calls needed. We're talking batch processing and careful prompting, not instant chat.</p>
<p>There is a tradeoff. Larger models mean slower inference, more heat, and louder fans. But here's the part that surprises most people.</p>
<p>Training an AI model requires massive, sustained compute. It's like running a marathon at sprint speed for weeks. But inference—just using the model to get an answer—doesn't need that same intensity. You're reading a huge matrix multiplication table, not calculating it from scratch.</p>
<p>That's why a laptop can do this. It's not performing miracles. It's just reading the answer, one compressed piece at a time.</p>
<hr />
<p><strong>Next in the series:</strong> <a href="#">The New Shape of AI Models →</a></p>
<p><strong>New to the jargon?</strong> Every term in this series has a plain-English explanation in the <a href="https://nightthoughts.hashnode.dev/plain-english-glossary-ai-terms-for-the-rest-of-us">glossary</a>.</p>
]]></content:encoded></item><item><title><![CDATA[Bigger Isn't Always Smarter]]></title><description><![CDATA[New to the jargon? Start with the plain-English glossary.

For years, the rule was simple: bigger AI meant smarter AI.
More parameters, more brainpower. Every new model was just the last one, scaled u]]></description><link>https://nightthoughts.me/bigger-isn-t-always-smarter</link><guid isPermaLink="true">https://nightthoughts.me/bigger-isn-t-always-smarter</guid><category><![CDATA[AI]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[local ai]]></category><category><![CDATA[technology]]></category><category><![CDATA[llm]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 04 Oct 2026 09:33:44 GMT</pubDate><content:encoded><![CDATA[<p><em>New to the jargon? Start with the <a href="https://nightthoughts.hashnode.dev/plain-english-glossary-ai-terms-for-the-rest-of-us">plain-English glossary</a>.</em></p>
<hr />
<p>For years, the rule was simple: bigger AI meant smarter AI.</p>
<p>More parameters, more brainpower. Every new model was just the last one, scaled up. A 70-billion-parameter model was smarter than a 7-billion one, the way a bigger library holds more books. And if you wanted to run one of those giants, you needed a data center.</p>
<p>That rule just broke.</p>
<p>A 70-billion-parameter model now runs on a 4GB graphics card — the kind you'd find in a mid-range gaming laptop. Slowly, yes. But it runs. Under the old rule, that shouldn't be possible.</p>
<p>To understand why it matters, it helps to see where the old rule came from — and why it was right for so long.</p>
<h2>The rule that worked</h2>
<p>In 2020, researchers noticed something remarkable. As they made AI models bigger, the models got better in a way that was almost predictable. Double the size, and performance improved by a measurable amount. Double it again, and it improved again.</p>
<p>They called this "scaling laws." And for a few years, scaling laws were the closest thing AI had to a recipe.</p>
<p>The recipe went something like this:</p>
<ul>
<li>More parameters: better.</li>
<li>More training data: better.</li>
<li>More compute: better.</li>
</ul>
<p>If you wanted a smarter model, you made a bigger one.</p>
<p>And it worked. Models went from millions of parameters to billions, then to hundreds of billions. Each generation outclassed the last. The rule seemed less like a pattern and more like a law of nature.</p>
<p>Let's be fair: the old rule wasn't wrong. It was incomplete. For most of a decade, scaling up really did work. The problem wasn't that the rule failed — it's that we mistook it for the whole story.</p>
<p>Because something else was happening underneath.</p>
<h2>The crack</h2>
<p>In 2022, a paper called <a href="https://arxiv.org/abs/2203.15556">Chinchilla</a> made a quiet but important point. Many of the biggest models weren't just big — they were undertrained. They had billions of parameters but hadn't seen enough data to use them well.</p>
<p>The implication was uncomfortable: a smaller model, trained on more data, could beat a bigger one trained on less.</p>
<p>At first, this was a footnote. Then it became a pattern. Small models started showing up on leaderboards next to giants. A well-trained 7-billion-parameter model could hold its own against one three times its size. In some cases, it beat them outright.</p>
<p>The old rule had a hole in it. And once you saw the hole, you couldn't unsee it.</p>
<h2>Why the rule broke</h2>
<p>Four things changed at roughly the same time. None of them alone explains the shift, but together they rewrote the rule.</p>
<p><strong>Better data.</strong> For years, "more data" meant "scrape more of the internet." Then researchers started curating it — filtering, cleaning, and in some cases generating high-quality training data on purpose. A model fed clean, dense, well-chosen examples learns faster than one fed a mountain of noise. Quality started to matter as much as quantity.</p>
<p><strong>Better teaching.</strong> Techniques like distillation let a small model learn from a big one. Instead of training from scratch, the small model copies the big model's answers and reasoning, the way an apprentice learns by watching a master. The result is a model a fraction of the size, with a surprising amount of the capability.</p>
<p><strong>Smarter designs.</strong> The biggest shift might be architectural. Mixture-of-experts models are built from many specialist sub-models. When a question comes in, only a few of them wake up to answer. On paper, the model is enormous. In practice, only a slice of it runs at any given moment. Huge on paper, cheap in practice.</p>
<p><strong>Better hardware.</strong> New chips, new number formats, and new software tricks let small machines hold more of the model at once. We'll dig into this in the next essay, but the short version is: the hardware caught up to the software's ambitions.</p>
<p><img src="https://datawrapper.dwcdn.net/xnb1K/full.png" alt="A scatter plot showing parameter count on the horizontal axis and benchmark performance on the vertical axis. Many small models appear above the trend line, and many larger models appear below it, illustrating that size alone does not predict capability." /></p>
<p><em>Sources: Official model cards, technical reports, and papers from Meta, Mistral AI, Alibaba Cloud, IBM, and DeepSeek (2023-2024). Scores are approximate and rounded.</em></p>
<p>Each of these forces pushed in the same direction: capability stopped being tied to raw size. A smaller model, designed and trained well, could do what used to require something much bigger.</p>
<h2>What "smart" actually means</h2>
<p>Here's the part that's easy to miss. "Smarter" was never a single number.</p>
<p>When we say a model is smart, we usually mean it's good at a specific set of tasks: answering questions, writing code, summarizing documents, reasoning through a math problem. But no model is good at everything. A model that writes beautiful prose might fail at arithmetic. A model that codes brilliantly might be hopeless at creative writing.</p>
<p>Benchmarks try to measure this, but they're imperfect. They get gamed. They get outdated. A model that tops one leaderboard can look mediocre on another.</p>
<p>So when we say "bigger isn't smarter," we're not saying size doesn't matter. We're saying size is one dimension of a model that has many dimensions — and that the other dimensions have gotten much more interesting.</p>
<h2>The new question</h2>
<p>If size isn't the measure, what is?</p>
<p>The honest answer is: it depends on what you need. For a chatbot that responds instantly, speed matters most. For a background process that runs overnight, speed matters less than cost. For a medical application, privacy might matter more than either.</p>
<p>The old question was "how big is it?" The new question is closer to: "how useful is it, how cheap is it, and how fast can it run — for what I'm actually doing?"</p>
<p>That question is harder to answer. But it's a much better one.</p>
<p>And it's the reason a 70-billion-parameter model can now live on a laptop.</p>
<p>That story — the how — is the next essay.</p>
]]></content:encoded></item><item><title><![CDATA[Plain English Glossary: AI Terms for the Rest of Us]]></title><description><![CDATA[The Basics



Term
What it actually means



Parameter
A dial the model tuned while learning. Billions of them. Think of them as the model's settings.


Model size (e.g., 70B)
How many dials the model]]></description><link>https://nightthoughts.me/plain-english-glossary-ai-terms-for-the-rest-of-us</link><guid isPermaLink="true">https://nightthoughts.me/plain-english-glossary-ai-terms-for-the-rest-of-us</guid><category><![CDATA[glossary]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[local ai]]></category><category><![CDATA[beginners guide]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 04 Oct 2026 02:21:12 GMT</pubDate><content:encoded><![CDATA[<h2>The Basics</h2>
<table>
<thead>
<tr>
<th>Term</th>
<th>What it actually means</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Parameter</strong></td>
<td>A dial the model tuned while learning. Billions of them. Think of them as the model's settings.</td>
</tr>
<tr>
<td><strong>Model size (e.g., 70B)</strong></td>
<td>How many dials the model has. 70B = 70 billion. Bigger isn't automatically smarter.</td>
</tr>
<tr>
<td><strong>Quantization</strong></td>
<td>Storing the model in lower resolution, like a compressed photo. Slightly blurrier, much smaller.</td>
</tr>
<tr>
<td><strong>FP16 / INT8 / INT4</strong></td>
<td>The "resolution" of the model. FP16 is high-res (2 bytes per parameter). INT4 is compressed (0.5 bytes each). Lower numbers mean smaller files but slightly lower quality.</td>
</tr>
<tr>
<td><strong>Inference</strong></td>
<td>Running the model to get an answer. The moment you hit "send."</td>
</tr>
<tr>
<td><strong>Training</strong></td>
<td>The long, expensive process of teaching the model. Happens once, before you ever use it.</td>
</tr>
<tr>
<td><strong>Tokens</strong></td>
<td>Chunks of text the model reads and writes. Roughly ¾ of a word each.</td>
</tr>
<tr>
<td><strong>Tokens per second</strong></td>
<td>How fast the AI writes. 1–2/sec is a crawl. 40+/sec feels instant.</td>
</tr>
<tr>
<td><strong>Time-to-first-token</strong></td>
<td>How long you wait before the AI starts replying. The "thinking" pause.</td>
</tr>
<tr>
<td><strong>VRAM</strong></td>
<td>Special memory on a graphics card. The model has to fit here to run fast.</td>
</tr>
<tr>
<td><strong>RAM</strong></td>
<td>Your computer's regular memory. More RAM = you can hold more of the model at once.</td>
</tr>
<tr>
<td><strong>KV cache</strong></td>
<td>The model's short-term memory of your conversation so far. It grows as you chat.</td>
</tr>
</tbody></table>
<hr />
<h2>The Runner-Ups</h2>
<table>
<thead>
<tr>
<th>Term</th>
<th>What it actually means</th>
</tr>
</thead>
<tbody><tr>
<td><strong>MoE (Mixture of Experts)</strong></td>
<td>A model made of specialist teams. Only a few wake up per question, so it's cheaper to run.</td>
</tr>
<tr>
<td><strong>Active vs. total parameters</strong></td>
<td>How many dials are <em>used</em> per answer vs. how many <em>exist</em>. A huge model can still be cheap.</td>
</tr>
<tr>
<td><strong>Dense model</strong></td>
<td>A model that uses all its dials for every answer. Simple but expensive.</td>
</tr>
<tr>
<td><strong>Sparse model</strong></td>
<td>A model that only uses part of itself per answer. Cheaper, often just as good.</td>
</tr>
<tr>
<td><strong>Pruning</strong></td>
<td>Trimming the model. Removing weights that aren't pulling their weight (setting them to zero) so the model takes up less space and runs faster.</td>
</tr>
<tr>
<td><strong>Distillation</strong></td>
<td>Teaching a small model to copy a big one. Like a student learning from a professor.</td>
</tr>
<tr>
<td><strong>Sparse attention</strong></td>
<td>A shortcut for long conversations. Instead of looking at every word it has ever seen, the model focuses only on the parts that matter right now.</td>
</tr>
<tr>
<td><strong>Speculative decoding</strong></td>
<td>Guessing ahead to speed up answers. If the guess is right, you save time.</td>
</tr>
<tr>
<td><strong>Test-time compute</strong></td>
<td>Letting the model think longer before answering. Smarter, but slower.</td>
</tr>
<tr>
<td><strong>NPU</strong></td>
<td>A chip built specifically for AI. Like a GPU, but for math instead of graphics.</td>
</tr>
<tr>
<td><strong>Layer streaming / offloading</strong></td>
<td>Keeping most of the model on disk, pulling in only what's needed, one piece at a time.</td>
</tr>
<tr>
<td><strong>NVFP4 / FP8</strong></td>
<td>Extra-compact number formats. Save memory, keep most of the quality.</td>
</tr>
<tr>
<td><strong>Local inference</strong></td>
<td>Running the model on your own machine. Private, offline, no subscription.</td>
</tr>
<tr>
<td><strong>Cloud inference</strong></td>
<td>Running the model on someone else's servers. Faster, stronger, but your data leaves home.</td>
</tr>
<tr>
<td><strong>Hybrid routing</strong></td>
<td>Small model handles easy questions; cloud handles the hard ones. The best of both.</td>
</tr>
</tbody></table>
<hr />
<h2>Going Deeper (Optional)</h2>
<p>These come up less often, but you might see them in the wild.</p>
<table>
<thead>
<tr>
<th>Term</th>
<th>What it actually means</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Fine-tuning</strong></td>
<td>Teaching an already-trained model a specific skill.</td>
</tr>
<tr>
<td><strong>RAG (Retrieval-Augmented Generation)</strong></td>
<td>Letting the model look things up before answering.</td>
</tr>
<tr>
<td><strong>Context window</strong></td>
<td>How much text the model can "hold in mind" at once.</td>
</tr>
<tr>
<td><strong>Benchmark</strong></td>
<td>A standard test used to compare models. Useful, but imperfect.</td>
</tr>
<tr>
<td><strong>Scaling laws</strong></td>
<td>The old rule that bigger models get predictably better. Now more nuanced.</td>
</tr>
<tr>
<td><strong>FLOPs</strong></td>
<td>A measure of raw computing work. Like counting how many math problems a chip can do.</td>
</tr>
<tr>
<td><strong>Unified memory</strong></td>
<td>A design where the CPU and GPU share the same pool of memory. Apple's approach.</td>
</tr>
</tbody></table>
<hr />
<p><em>This glossary is part of the <a href="https://nightthoughts.hashnode.dev/series/big-models-small-machines"><strong>Big Models, Small Machines</strong></a> series.</em></p>
]]></content:encoded></item><item><title><![CDATA[The Qwen Uncensored Model: How "Brain Surgery" Removed AI Safety Rules
]]></title><description><![CDATA[Imagine downloading an AI onto your laptop. It is smart, fast, and can answer almost any question. But it is missing one big thing: the word "no."
This is the reality of a new model called Qwen 3.8 27]]></description><link>https://nightthoughts.me/the-qwen-uncensored-model-how-brain-surgery-removed-ai-safety-rules</link><guid isPermaLink="true">https://nightthoughts.me/the-qwen-uncensored-model-how-brain-surgery-removed-ai-safety-rules</guid><category><![CDATA[AI]]></category><category><![CDATA[AI Safety]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[llm]]></category><category><![CDATA[abliteration]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sat, 03 Oct 2026 15:04:31 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/002b1dcc-a0f4-4217-b586-f8d87299f469.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine downloading an AI onto your laptop. It is smart, fast, and can answer almost any question. But it is missing one big thing: the word "no."</p>
<p>This is the reality of a new model called <strong>Qwen 3.8 27B Uncensored</strong>. Normally, AI companies build safety rules into their models. If you ask a standard AI how to make a dangerous weapon or write a computer virus, it will refuse. It will say, "I can't help with that."</p>
<p>But this version of Qwen doesn't do that. It will answer. It will not warn you. It will just give you the information you asked for.</p>
<p>This is not a mistake or a bug. It was done on purpose using a method called <strong>abliteration</strong>. It is a way to permanently remove an AI's ability to refuse a request. And because the code and the model are now public, the genie is out of the bottle.</p>
<h3>The "Brain Surgery" Explained</h3>
<p>To understand how this happened, we have to look at how AI makes decisions.</p>
<p>Think of an AI model like a giant, complex brain made of millions of connections. When you ask it a question, a signal travels through those connections. Normally, AI companies train their models to have a safety feature. We can call this the <strong>"refusal button."</strong></p>
<p>If you ask a standard AI a dangerous question, the signal hits that button. The AI stops, and out comes a refusal: "I can't help with that."</p>
<p>Abliteration is a mathematical trick that finds that exact refusal button and <strong>deletes it.</strong></p>
<p>It is not just telling the AI to ignore its rules. The button is physically gone. The math that creates the refusal is cut out of the model's code. The AI literally loses the ability to say no.</p>
<p>The most surprising part is that this doesn't require retraining the AI. It's not like teaching a dog a new trick. It's more like taking a car and removing the brakes. The car still drives perfectly fine. It's still fast and powerful. But it simply cannot stop anymore.</p>
<p>This is exactly what happened to the Qwen model. The creators took the original, safety-trained Qwen 3.8 27B and performed this "brain surgery." The result is a model that is just as smart and just as capable as the original, but completely unable to refuse a request.</p>
<h3>The Qwen Model: Anyone Can Run It</h3>
<p>So, what exactly is this Qwen model?</p>
<p>Qwen is a family of AI models created by the Chinese tech company Alibaba. They released the original version as "open-source," which means they shared the underlying code with the public. Anyone can download it, use it, or change it.</p>
<p>This is where the problem starts. Because the code is public, anyone with a decent computer and some technical knowledge can take the Qwen model and perform the "brain surgery" we just talked about. That is exactly what happened here.</p>
<p>The result is the <strong>Qwen 3.8 27B Uncensored</strong> model.</p>
<p>What makes this so concerning is how easy it is to run. You do not need a giant data center or a supercomputer. According to the model's documentation, it can run on:</p>
<ul>
<li><p><strong>A Mac:</strong> You need about 13.5 GB of free memory. If you have a newer Mac with an Apple Silicon chip, it runs even faster.</p>
</li>
<li><p><strong>A Windows PC:</strong> You need a good graphics card (GPU) with about 16.8 GB of memory.</p>
</li>
</ul>
<p>This means a normal person can download this model onto their laptop. They can run it completely offline. No internet connection is needed. No company is watching what they ask. It is a powerful, completely unrestricted AI, sitting right on someone's desk.</p>
<h3>The Genie Out of the Bottle</h3>
<p>This brings us to the biggest problem of all: we cannot stop it.</p>
<p>Once the mathematical method for abliteration was shared online, it was over. You cannot delete a math formula from the internet. You cannot un-share code. The method is out there, and it works.</p>
<p>We call this the "genie out of the bottle" problem.</p>
<p>In the past, if a company like OpenAI or Google made a dangerous AI, they could just shut down their servers. They could stop people from using it. But you cannot shut down a file that has been downloaded to a million different laptops around the world.</p>
<p>There is no central server to ban. There is no company to sue. The model is just a file, and it is free.</p>
<p>The Qwen Uncensored model proves that AI safety rules are not permanent. They are just a thin layer that can be peeled away. And now that everyone knows how to peel it away, they can do it to almost any other open-source AI model in the future.</p>
<h3>The Bigger Picture</h3>
<p>It is important to understand why people do this.</p>
<p>Some people argue that standard AI models are too strict. They refuse to help with harmless things, like writing a dark fiction story or asking about medical symptoms. They want an AI that is free from corporate rules. They believe in total privacy and total freedom.</p>
<p>But there is a serious downside.</p>
<p>A tool that cannot say "no" is a danger to everyone. A bad actor can use this Qwen model to write convincing scams, create computer viruses, or plan harmful acts. There are no guardrails. The AI will not warn the user. It will not alert the authorities. It will simply help.</p>
<p>We are now living in a world where powerful AI has no safety net. The genie is out of the bottle, and it is not going back in. The question is no longer whether we can stop it. The real question is: how do we live in a world where anyone can access a brilliant, obedient, and completely amoral digital assistant?</p>
]]></content:encoded></item><item><title><![CDATA[Why AI Agents Need Governance—and How Qodo Provides It]]></title><description><![CDATA[AI agents can now write code. They can open pull requests, fix bugs, and add features at a speed humans cannot match. They do not get tired. That sounds like a big win for software teams.
But speed al]]></description><link>https://nightthoughts.me/why-ai-agents-need-governance-and-how-qodo-provides-it</link><guid isPermaLink="true">https://nightthoughts.me/why-ai-agents-need-governance-and-how-qodo-provides-it</guid><category><![CDATA[AI]]></category><category><![CDATA[Governance]]></category><category><![CDATA[qodo]]></category><category><![CDATA[code review]]></category><category><![CDATA[Devops]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sat, 03 Oct 2026 02:20:50 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/11f93c22-1c80-4473-afce-1a4b69fccb01.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI agents can now write code. They can open pull requests, fix bugs, and add features at a speed humans cannot match. They do not get tired. That sounds like a big win for software teams.</p>
<p>But speed alone is not enough. If many agents write code at the same time, teams can lose control. Code may not follow company rules. Reviews may pile up. No one may see the full picture. A small change in one place can break something far away. A missed review can become a security problem. This is why control and governance are necessary.</p>
<p>Control means setting clear rules. Governance means making sure those rules are followed and that people can see what is happening. With AI agents, this is hard because work is spread across many repositories and pull requests. Teams need a way to guide agents, review their work, and understand the impact of their changes.</p>
<h2>What Is Qodo?</h2>
<p>Qodo is a platform that provides a quality and governance layer for software built with AI agents. The company was founded in 2022 by Itamar Friedman and Dedy Kredo. It was originally called CodiumAI and rebranded to Qodo in 2024 as the platform grew beyond just test generation. The founders have deep technical backgrounds: Friedman previously founded Visualead, which was acquired by Alibaba, and led machine vision work at Alibaba. [1]</p>
<h2>The Team and Backing Behind Qodo</h2>
<p>Qodo is backed by serious investors and advisors. The company has raised a total of $120 million, including a $70 million Series B round in March 2026. The round was led by Qumra Capital, with participation from Square Peg, Susa Ventures, TLV Partners, Vine Ventures, and others. Individual investors include Peter Welinder, VP of Product at OpenAI, and Clara Shih, VP of AI at Meta. [2]</p>
<p>The advisory board includes leaders from OpenAI, Meta, Shopify, and Snyk. This matters because it shows that people who understand both AI and enterprise software believe in what Qodo is building. The company is headquartered in New York. [3]</p>
<h2>Why Qodo Is Promising</h2>
<p>Qodo is not just another code review tool. It focuses on a problem that most tools ignore: understanding how a code change fits into the <em>whole system</em>. Most AI review tools look at what changed in a single pull request. Qodo looks at how that change affects the entire codebase, considering organizational standards, historical decisions, and risk tolerance.</p>
<p>The company’s CEO, Itamar Friedman, explained the thinking behind this: “Code generation companies are largely built around LLMs. But for code quality and governance, LLMs alone aren’t enough. Quality is subjective. It depends on organizational standards, past decisions, and tribal knowledge.” [4] This insight is the core of why Qodo exists.</p>
<p>The results speak for themselves. Qodo has an F1 score of 64.3% on the Code Review Bench, catching real problems at nearly twice the rate of others, including Claude. It is also ranked #1 by Gartner for code understanding in the Critical Capabilities for AI Assistants Report. Enterprises use it to catch an average of 800 bugs per month. [5]</p>
<h2>How Qodo 3.0 Solves the Governance Problem</h2>
<p>Qodo 3.0 is the latest version of the platform, released in October 2026. It adds new features designed specifically for governing AI-generated code at scale. Here is how it works:</p>
<ul>
<li><p><strong>PR Triage</strong> groups related pull requests across repositories into “work packages.” It scores them by impact, priority, and SLA. A reviewer can claim a package, so the same work is not reviewed twice. This directly addresses the review bottleneck that AI-generated code creates. [6]</p>
</li>
<li><p><strong>Agentic Toolbox</strong> lets coding agents use Qodo’s rules and codebase context <em>while they write</em>. This means company standards are applied before a pull request is even opened. The rules come from the organization’s own repositories, so agents work within the right guardrails from the start. [6]</p>
</li>
<li><p><strong>Software Map</strong> shows how repositories and services connect. It automatically maps dependencies, service relationships, and contracts. It also calculates the “blast radius” of a proposed change, so teams can see what might break. Quality issues appear as a heat map, making technical debt visible. [6]</p>
</li>
<li><p><strong>Analytics Dashboard</strong> shows how the team responds to Qodo findings. It tracks impact (how findings lead to accepted changes), compares activity by repository, and analyzes responses by PR author. This gives managers a clear picture of where quality is improving and where it is not. [6]</p>
</li>
<li><p><strong>Wisdom Base</strong> is the knowledge layer underneath everything. It keeps a continuously updated understanding of the codebase, standards, architecture, and pull request history. This is what makes the other features smart. It learns from the organization’s own patterns rather than applying generic rules. [6]</p>
</li>
<li><p><strong>Enterprise deployment</strong> options make adoption easier. There is an onboarding wizard, review presets tuned to different goals (catching more issues vs. reducing noise), rules import from specific repositories, and support for air-gapped or on-premises deployment. [6]</p>
</li>
</ul>
<h2>Licensing</h2>
<p>Qodo uses an <strong>open-core model</strong>. The core pull request review engine, known as PR-Agent, is open source. It was originally launched under the <strong>Apache 2.0 license</strong> in 2023. In April 2026, Qodo transferred stewardship of PR-Agent to a community-owned GitHub organization and restored the Apache 2.0 license, moving away from any restrictive terms. This means anyone can use, modify, and distribute the PR-Agent engine freely, and teams can self-host it by supplying their own LLM API keys and compute. [7]</p>
<p>The hosted <strong>Qodo Merge</strong> platform, along with the enterprise governance features introduced in Qodo 3.0, is <strong>closed and proprietary</strong>. Pricing follows a per-seat subscription model: a free Developer tier provides 30 PR reviews and 250 IDE/CLI credits per month, the Teams tier costs $30 per active user per month (billed annually), and Enterprise is custom-negotiated with options for self-hosted or on-premises deployment. This hybrid approach gives developers an open-source foundation to build on while providing enterprises with a commercial product that includes governance, analytics, and support. [8]</p>
<h2>Why Choose Qodo?</h2>
<p>The simple answer is this: Qodo understands that generating code and governing code are different problems. Most tools focus on generation. Qodo focuses on verification, quality, and governance.</p>
<p>Its edge comes from three things: deep context (it understands the whole system, not just one file), enterprise focus (it is built for large organizations with complex codebases), and a team with the right experience (founders who have built and sold companies, advisors from OpenAI and Meta, and investors who understand the space).</p>
<p>AI agents can help us build software faster. But without control and governance, speed can create risk. Qodo 3.0 aims to provide that control. It lets agents work quickly while humans stay in charge. That is the balance modern software teams need.</p>
<h2>References</h2>
<ol>
<li><p>Qodo. “Beyond Intelligence: Qodo’s $70M Series B and the Shift to Artificial Wisdom.” Qodo Blog, March 30, 2026. <a href="https://www.qodo.ai/blog/qodo-70m-series-b-shift-to-artificial-wisdom/">https://www.qodo.ai/blog/qodo-70m-series-b-shift-to-artificial-wisdom/</a></p>
</li>
<li><p>Qodo. “We’re on a mission to make code integrity simple.” Qodo About Page. <a href="https://www.qodo.ai/about/">https://www.qodo.ai/about/</a></p>
</li>
<li><p>Qodo. “Qodo | AI Agents for Code, Review &amp; Workflows.” Qodo Homepage. <a href="https://www.qodo.ai/">https://www.qodo.ai/</a></p>
</li>
<li><p>Qodo. “Critical Capabilities Report.” Qodo Reports, November 21, 2025. <a href="https://www.qodo.ai/reports/gartner-critical-capabilities-ai-code-assistance-2025/">https://www.qodo.ai/reports/gartner-critical-capabilities-ai-code-assistance-2025/</a></p>
</li>
<li><p>Qodo. “What’s new - Qodo Documentation.” Qodo Docs, October 1, 2026. <a href="https://docs.qodo.ai/whats-new">https://docs.qodo.ai/whats-new</a></p>
</li>
<li><p>Qodo. “Qodo 3.0 Puts Governance at the Center of Agentic Code.” Futurum Group, October 2, 2026. <a href="https://futurumgroup.com/qodo-3-0-puts-governance-at-the-center-of-agentic-code/">https://futurumgroup.com/qodo-3-0-puts-governance-at-the-center-of-agentic-code/</a></p>
</li>
<li><p>Qodo. “Qodo Is Handing PR-Agent Over to the Community.” Qodo Blog, April 23, 2026. <a href="https://www.qodo.ai/blog/qodo-is-handing-pr-agent-over-to-the-community/">https://www.qodo.ai/blog/qodo-is-handing-pr-agent-over-to-the-community/</a></p>
</li>
<li><p>API Evangelist. “Qodo Plans and Pricing.” API Commons, June 21, 2026. <a href="https://raw.githubusercontent.com/api-evangelist/qodo/refs/heads/main/plans/qodo-plans-pricing.yml">https://raw.githubusercontent.com/api-evangelist/qodo/refs/heads/main/plans/qodo-plans-pricing.yml</a></p>
</li>
<li><p>Qodo. “Qodo Raises $70M in Series B Funding.” FinSMEs, March 30, 2026. <a href="https://www.finsmes.com/2026/03/qodo-raises-70m-in-series-b-funding.html">https://www.finsmes.com/2026/03/qodo-raises-70m-in-series-b-funding.html</a></p>
</li>
<li><p>Qodo. “Qodo Ranked #1 AI Code Review Tool in Martian’s Code Review Benchmark.” Qodo Blog, March 15, 2026. <a href="https://www.qodo.ai/blog/qodo-ranked-1-ai-code-review-tool-in-martians-code-review-benchmark/">https://www.qodo.ai/blog/qodo-ranked-1-ai-code-review-tool-in-martians-code-review-benchmark/</a></p>
</li>
</ol>
]]></content:encoded></item><item><title><![CDATA[The Sovereignty Illusion: What Companies Get Wrong About Cloud Independence]]></title><description><![CDATA[This is the third chapter of Built to Leave, a series on cloud dependency, sovereignty, and the architecture that makes exit possible. The series follows the problem from its origins — the build-vs-bu]]></description><link>https://nightthoughts.me/the-sovereignty-illusion-what-companies-get-wrong-about-cloud-independence</link><guid isPermaLink="true">https://nightthoughts.me/the-sovereignty-illusion-what-companies-get-wrong-about-cloud-independence</guid><category><![CDATA[Sovereignty]]></category><category><![CDATA[Cloud Computing]]></category><category><![CDATA[data]]></category><category><![CDATA[finance]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Fri, 02 Oct 2026 15:06:09 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/9db712eb-14bc-4f6d-bfd3-6830e6342c6b.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is the third chapter of <strong>Built to Leave</strong>, a series on cloud dependency, sovereignty, and the architecture that makes exit possible. The series follows the problem from its origins — the build-vs-buy trade-off that predates AI — through the structural costs of cloud dependency, the work required to regain control, and the architectural shift that AI workloads are now forcing. Full table of contents: <a href="https://nightthoughts.hashnode.dev/build-to-leave">nightthoughts.hashnode.dev/build-to-leave</a></em></p>
<hr />
<p>In the first chapter of this series, I looked at the oldest trade-off in enterprise technology: build it yourself, or buy it from someone else. In the second, I looked at how two forces—defensive security and offensive competitiveness—are pushing financial institutions toward the same infrastructure at the same time. This chapter looks at the foundation both of those forces run on, and the uncomfortable question underneath it: <strong>whose cloud is it, really?</strong></p>
<p>There's a phrase that captures the core misunderstanding I want to unpack here: <em>"Your data, their laws."</em> It sounds provocative, maybe even alarmist. But for a European company storing data with a US-headquartered cloud provider, it is a literal description of the legal reality. The US CLOUD Act and FISA allow US authorities to compel US-incorporated companies to hand over data—regardless of where that data is physically stored. A server in Frankfurt does not put that data beyond the reach of US law if the company operating it is subject to US jurisdiction.</p>
<p>That's the sovereignty illusion. For years, "sovereign cloud" has been sold to European companies as a matter of data residency: keep the bits inside EU borders, and you've solved the problem. But jurisdiction follows the company, not the server. And as long as the parent company answers to another government's laws, "complete sovereignty" remains out of reach.</p>
<p>This matters most for Europe, where the gap between ambition and reality is widest. US hyperscalers control over 70% of the European cloud market. European providers hold around 15%, a share that has barely moved despite years of investment and political will. DORA has given European supervisors new powers over critical third-party providers, but oversight is not the same as alternatives. And the concentration risk isn't theoretical—the October 2025 AWS US-East-1 outage took down European banks, brokerages, and payment systems, a preview of what's at stake when the infrastructure you don't control fails.</p>
<p>For US or Chinese companies, this is rarely a live question—they operate under the same jurisdiction as their dominant cloud providers. For European companies, it is unavoidable. So what does "independence" actually mean when both your security posture and your competitive AI capability depend on infrastructure owned by someone else, operating under someone else's laws? That's the question this chapter sits with. It doesn't have a tidy answer—but the first step is naming the problem clearly: <strong>your data, their laws.</strong> Everything else follows from there.</p>
<hr />
<h2>Why Public Cloud Is Non-Negotiable</h2>
<p>Before getting into the risks, it's worth being honest about why European companies can't simply walk away from public cloud. The dependency isn't a failure of strategy; it's a rational response to three converging pressures.</p>
<p><strong>AI requires scale.</strong> Generative AI has deepened dependency across three layers: specialized chips (Nvidia-dominated), foundation models (GPT, Gemini, Claude), and integrated AI platforms from hyperscalers. Building that stack in-house is not a realistic option for most companies. The compute requirements alone—training runs costing tens of millions, inference at scale requiring global distribution—put it out of reach for any institution whose core business is finance, not infrastructure.</p>
<p><strong>Resilience demands it.</strong> Rising climate and geopolitical risks are pushing infrastructure resilience requirements toward multi-region, geographically diversified setups almost as a baseline expectation now, not a best practice for the ambitious. One global analysis of nearly 9,000 data centers found that a meaningful share already sit in locations facing high or moderate physical climate risk. A separate analysis put the share of capacity exposed to acute climate hazards—flooding, extreme wind, wildfire—at close to four in five facilities worldwide. Meeting that bar in-house means duplicating infrastructure across regions, at a cost and complexity that grows with every added region.</p>
<p><strong>Cost efficiency is real.</strong> For years, US hyperscalers offered European companies unmatched scalability and cost efficiency. The pay-as-you-go model converts large fixed investments into variable costs. For institutions under margin pressure, that flexibility is not a luxury—it's a competitive necessity.</p>
<p><strong>Regulatory pressure cuts both ways.</strong> DORA requires robust operational resilience, but it also acknowledges that cloud concentration is a systemic risk. The regulation gives supervisors new powers over critical third-party providers, but oversight is not the same as alternatives. A covered entity using a designated provider gets some assurance that the provider is being supervised, but the covered entity remains accountable for its own risk management.</p>
<p>The point is this: public cloud isn't a choice European companies can undo. The question is what that dependency actually costs—and whether the price is being measured honestly.</p>
<hr />
<h2>The Risks: What Companies Are Actually Paying For</h2>
<h3>The Jurisdiction Trap</h3>
<p>The US CLOUD Act and FISA allow US authorities to compel US-incorporated companies to hand over data regardless of where it's stored. This creates a direct conflict with GDPR and the EU's data protection framework. A server in Frankfurt is legally an American server. Data residency only answers where data sits, not who can be compelled to hand it over.</p>
<p>This isn't a theoretical concern. The CLOUD Act was designed precisely to resolve conflicts between US law enforcement demands and foreign data protection laws—and it resolves them in favor of US access. For a European company, that means the legal protections you've built around data residency can be overridden by a legal order you have no standing to contest.</p>
<h3>Concentration Risk</h3>
<p>According to the Dutch Authority for the Financial Markets (AFM) and De Nederlandsche Bank (DNB), the Dutch financial sector faces increasing systemic risks stemming from its growing reliance on a limited number of non-European IT service providers. The regulators caution that this dependency amplifies the risk of concentration and systemic disruption, where a failure at a single provider could impact large segments of the financial sector (<a href="https://www.afm.nl/en/sector/actueel/2025/okt/pb-digitale-autonomie">AFM &amp; DNB, 2025</a>).</p>
<p>More than 30% of significant EU institutions' total outsourcing budgets flow to just 10 providers. In 2025, 29% of major ICT incidents reported across the EU financial sector originated from third-party provider failures. The October 2025 AWS US-East-1 outage took down banks, brokerages, and payment systems across the continent—a preview of what happens when infrastructure you don't control fails.</p>
<p>The concentration problem isn't just about outages. It's about the fact that a handful of providers now sit underneath the operational resilience of the entire European financial system. When many institutions lean on the same few providers, a single failure—technical, legal, or geopolitical—ripples across all of them at once.</p>
<h3>Geopolitical and Regulatory Exposure</h3>
<p>Model availability can be restricted through US sanctions or export controls. Sensitive data and metadata may be exposed to third-country access. Access to cloud services, updates, or AI models could be restricted in a crisis scenario.</p>
<p>This is the risk that's hardest to price because it's the hardest to imagine. But it's not hypothetical. Export controls have already been used to restrict access to advanced chips and AI models. Sanctions can sever access to services overnight. Data localization mandates can force abrupt architectural changes. For a European company that has built its AI strategy on a US hyperscaler's platform, a geopolitical rupture isn't just a business continuity problem—it's an existential one.</p>
<h3>The Cost of Sovereignty</h3>
<p>According to Eric Bierry, CEO of SBS and Deputy CEO of 74 Software, when CDC, the French public bank, migrated a regulatory reporting service from AWS to France's sovereign SecNumCloud infrastructure, <strong>the cost more than tripled</strong>—even though the underlying product and service stayed exactly the same (<a href="https://sbs-software.com/insights/artificial-intelligence-data/podcast-ai-in-banking-why-100-sovereign-is-a-myth/">SBS Software, 2026</a>).</p>
<p>True 100% sovereignty isn't achievable anyway. Even domestic infrastructure relies on non-European chips and software somewhere in the stack (<a href="https://sbs-software.com/insights/artificial-intelligence-data/podcast-ai-in-banking-why-100-sovereign-is-a-myth/">SBS Software, 2026</a>). The question isn't whether you can achieve perfect sovereignty—you can't—but whether you're making deliberate trade-offs between cost, control, and risk.</p>
<h3>The Compliance Burden</h3>
<p>The US CLOUD Act and FISA enable government access to data physically held in the EU where a US nexus exists. This creates a direct conflict with GDPR and the EU's data protection framework. Companies are caught between two legal regimes that make incompatible demands: US law says hand it over; EU law says you can't.</p>
<p>That conflict doesn't resolve itself through better contracts or more careful data mapping. It's structural. And it's the reason "sovereign cloud" keeps coming up in board conversations that would rather be about something else.</p>
<hr />
<h2>What Companies Can Do: Recommendations</h2>
<p>The risks are real, but they're not unmanageable. What follows is a set of practical recommendations drawn from regulators, industry bodies, and real-world migrations—not a solution set, but a starting point for institutions that need to act while the strategic questions are still being worked out.</p>
<p><strong>1. Map concentration risk by service, not just provider.</strong> According to AFME's 2021 paper on cloud computing, banks need to proactively architect for greater resilience by mapping dependencies between services and geographies—identifying, for example, where two different services share a single point of failure, or how an outage in one region may affect the underlying cloud service provider's control plane (<a href="https://www.afme.eu/Portals/0/DispatchFeaturedImages/AFME_CloudComputing2021_06-2.pdf">AFME, 2021</a>). Concentration risk is an aggregate of your underlying threat scenarios. You need to evaluate each service dependency individually, not just count how many hyperscalers you use.</p>
<p><strong>2. Adopt a hybrid and multi-cloud strategy.</strong> AFME's recommendations emphasize that banks should maintain an inventory of cloud arrangements and take a risk-based approach using a range of approaches, including multi-cloud and portability (<a href="https://www.afme.eu/Portals/0/DispatchFeaturedImages/AFME_CloudComputing2021_06-2.pdf">AFME, 2021</a>). In practice: sensitive and regulated data stays within EU-based sovereign clouds, critical systems remain on-premises for maximum control, and less sensitive workloads continue on US hyperscale platforms.</p>
<p><strong>3. Build a placement and classification logic.</strong> DORA requires financial entities to maintain an ICT third-party risk strategy and policy for ICT services supporting critical or important functions (<a href="https://www.eba.europa.eu/publications-and-media/press-releases/european-supervisory-authorities-designate-critical-ict-third-party-providers-under-digital">DORA</a>). Classify applications, data, and processes by criticality, regulatory sensitivity, data confidentiality, substitutability, and dependency risk. This determines where each workload should be deployed. EU-based operations may be appropriate for critical services, but they do not guarantee sovereignty on their own—you must also assess control rights, support models, access paths, encryption concepts, and contractual exit options.</p>
<p><strong>4. Establish and test exit strategies before go-live.</strong> DORA Article 28(8) requires supervised entities to develop transition plans enabling them to remove contracted ICT services and data from third-party providers and securely transfer them to alternatives or in-house (<a href="https://www.eba.europa.eu/publications-and-media/press-releases/european-supervisory-authorities-designate-critical-ict-third-party-providers-under-digital">DORA</a>). According to the ECB's supervisory guide, exit plans must be realistic, viable, based on plausible scenarios, and include a planned execution timeline compatible with contractual exit clauses (<a href="https://www.bankingsupervision.europa.eu/ecb/pub/pdf/ssm.supervisory_guides202507.pt.pdf">ECB, 2025</a>). Their feasibility should be tested with independent verification.</p>
<p><strong>5. Maintain independent control over encryption keys.</strong> According to the European Banking Authority's outsourcing guidelines (EBA/GL/2019/02), financial institutions must retain effective control over encryption keys protecting sensitive data processed or stored by third-party service providers (<a href="https://www.kiteworks.com/regulatory-compliance/eba-encryption-key-control-guidelines/">EBA, 2019</a>). Without direct key control, institutions cannot demonstrate data sovereignty, execute exit strategies, or guarantee recovery during vendor failures or disputes (<a href="https://www.kiteworks.com/regulatory-compliance/eba-encryption-key-control-guidelines/">EBA, 2019</a>). The EBA treats encryption key control as a fundamental prerequisite for maintaining operational resilience and data sovereignty when delegating functions to external service providers.</p>
<p><strong>6. Negotiate collectively.</strong> The AFM and DNB have explicitly warned that widespread reliance on the same providers and infrastructures has led to concentration and systemic risks, and they urge institutions to prepare for disruptive scenarios by collaborating with IT vendors, authorities, and peers (<a href="https://www.afm.nl/en/sector/actueel/2025/okt/pb-digitale-autonomie">AFM &amp; DNB, 2025</a>). A single company negotiating alone with Amazon or Microsoft has "hardly anything to say"—a few hundred together do.</p>
<p><strong>7. Invest in European sovereign providers where it makes strategic sense.</strong> The ECB itself chose OVHcloud to provide sovereign cloud services for the digital euro project, with infrastructure operated entirely within the European Union (<a href="https://corporate.ovhcloud.com/en-ca/newsroom/news/ovhcloud-digital-euro-ecb/">OVHcloud, 2026</a>). De Nederlandsche Bank (DNB) signed a contract with Stackit, the cloud platform owned by Germany's Schwarz Group, to reduce its dependence on American cloud companies (<a href="https://www.techzine.eu/news/infrastructure/140634/dutch-central-bank-chooses-lidl-for-european-cloud/">Techzine, 2026</a>). Smaller institutions have fewer resources but also lower barriers to adopting European providers—and regulators reward early movers, so an apparent weakness can become a strategic strength.</p>
<p><strong>8. Treat digital sovereignty as a board-level concentration risk topic.</strong> The AFM and DNB emphasize that reducing digital dependence is a long-term challenge requiring coordinated European solutions, and they call for the development of robust European alternatives to non-EU IT providers (<a href="https://www.afm.nl/en/sector/actueel/2025/okt/pb-digitale-autonomie">AFM &amp; DNB, 2025</a>). It connects operational resilience, outsourcing governance, technology strategy, data protection, AI adoption, and geopolitical risk management. The question is no longer whether companies should use global technology platforms, but whether they can use them without losing transparency, decision rights, portability, and credible exit options for their most critical services.</p>
<hr />
<h2>The Question That Doesn't Go Away</h2>
<p>None of these recommendations solve the underlying problem. They manage it. They buy time. They reduce exposure at the margins. But they don't change the fundamental reality: European companies depend on infrastructure owned by companies that answer to another government's laws.</p>
<p>That dependency isn't going away. Public cloud is too central to AI competitiveness, too embedded in resilience planning, too cost-effective to abandon. The question isn't whether to use it—it's whether the terms of that use can be renegotiated in a way that gives European companies meaningful control over their own security, their own data, and their own competitive future.</p>
<p>The next chapter looks at what solutions actually exist—not just risk management, but structural alternatives. What would a genuinely sovereign cloud stack look like? What's being built, what's working, and what's still missing? And what would it take for European companies to have real options, not just better contracts?</p>
<p>That's where we go next.</p>
<hr />
<h3>References</h3>
<ul>
<li><p>PIFS International, <a href="https://www.pifsinternational.org/cloud-adoption-in-the-financial-sector-and-concentration-risk/">"Cloud Adoption in the Financial Sector and Concentration Risk"</a></p>
</li>
<li><p>Office of the Superintendent of Financial Institutions (OSFI), <a href="https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/third-party-risk-management-guideline">"Third-Party Risk Management Guideline"</a></p>
</li>
<li><p>Bristows, <a href="https://inquisitiveminds.bristows.com/post/102lqkb/aws-us-east-1-incident-regulators-concentrate-on-concentration-risk">"AWS US-EAST-1 incident: regulators concentrate on concentration risk"</a></p>
</li>
<li><p>calQrisk, <a href="https://www.calqrisk.com/resources/insights/outsourcing-and-third-party-risk-management-for-financial-firms">"Outsourcing and Third-Party Risk Management for Financial Firms"</a></p>
</li>
<li><p>FINRA, <a href="https://www.finra.org/rules-guidance/guidance/reports/2025-finra-annual-regulatory-oversight-report/third-party-risk">"Third-Party Risk Landscape," 2025 Annual Regulatory Oversight Report</a></p>
</li>
<li><p>Data Center Dynamics, <a href="https://www.datacenterdynamics.com/en/news/climate-threats-to-data-centers-set-to-surge-report/">"Climate threats to data centers set to surge"</a></p>
</li>
<li><p>XDI, <a href="https://xdi.systems/news/global-data-centres-face-rising-climate-risks-xdi-report-warns-landmark-analysis-of-nearly-9000-sites-reveals-escalating-threat-to-digital-infrastructure/">"Global data centres face rising climate risks"</a></p>
</li>
<li><p>European Supervisory Authorities, DORA CTPP Designation List (November 2025) — <a href="https://www.eba.europa.eu/publications-and-media/press-releases/european-supervisory-authorities-designate-critical-ict-third-party-providers-under-digital">https://www.eba.europa.eu/publications-and-media/press-releases/european-supervisory-authorities-designate-critical-ict-third-party-providers-under-digital</a></p>
</li>
<li><p>ECB, Supervisory Guide on Cloud Outsourcing (July 2025) — <a href="https://www.bankingsupervision.europa.eu/ecb/pub/pdf/ssm.supervisory_guides202507.pt.pdf">https://www.bankingsupervision.europa.eu/ecb/pub/pdf/ssm.supervisory_guides202507.pt.pdf</a></p>
</li>
<li><p>EBA, Outsourcing Guidelines (EBA/GL/2019/02) — <a href="https://www.kiteworks.com/regulatory-compliance/eba-encryption-key-control-guidelines/">https://www.kiteworks.com/regulatory-compliance/eba-encryption-key-control-guidelines/</a></p>
</li>
<li><p>AFM &amp; DNB, "AFM and DNB warn of systemic risks in the financial sector from digital dependence" (October 2025) — <a href="https://www.afm.nl/en/sector/actueel/2025/okt/pb-digitale-autonomie">https://www.afm.nl/en/sector/actueel/2025/okt/pb-digitale-autonomie</a></p>
</li>
<li><p>OVHcloud, "OVHcloud to provide sovereign cloud services for the ECB digital euro" (March 2026) — <a href="https://corporate.ovhcloud.com/en-ca/newsroom/news/ovhcloud-digital-euro-ecb/">https://corporate.ovhcloud.com/en-ca/newsroom/news/ovhcloud-digital-euro-ecb/</a></p>
</li>
<li><p>Techzine, "Dutch central bank chooses Lidl for European Cloud" (April 2026) — <a href="https://www.techzine.eu/news/infrastructure/140634/dutch-central-bank-chooses-lidl-for-european-cloud/">https://www.techzine.eu/news/infrastructure/140634/dutch-central-bank-chooses-lidl-for-european-cloud/</a></p>
</li>
<li><p>SBS Software, "AI in banking | Part 3: Why '100% sovereign' is a myth" (September 2026) — <a href="https://sbs-software.com/insights/artificial-intelligence-data/podcast-ai-in-banking-why-100-sovereign-is-a-myth/">https://sbs-software.com/insights/artificial-intelligence-data/podcast-ai-in-banking-why-100-sovereign-is-a-myth/</a></p>
</li>
<li><p>AFME, "Cloud Computing in Capital Markets" (June 2021) — <a href="https://www.afme.eu/Portals/0/DispatchFeaturedImages/AFME_CloudComputing2021_06-2.pdf">https://www.afme.eu/Portals/0/DispatchFeaturedImages/AFME_CloudComputing2021_06-2.pdf</a></p>
</li>
<li><p>FISA (Foreign Intelligence Surveillance Act), including Section 702 provisions — <a href="https://www.law.cornell.edu/wex/foreign_intelligence_surveillance_act">https://www.law.cornell.edu/wex/foreign_intelligence_surveillance_act</a></p>
</li>
<li><p>US CLOUD Act — <a href="https://www.congress.gov/bill/115th-congress/senate-bill/2383">https://www.congress.gov/bill/115th-congress/senate-bill/2383</a></p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[Two Races at Once: Why AI Is Both a Security Emergency and a Competitive One]]></title><description><![CDATA[This is the second chapter of Built to Leave, a series on cloud dependency, sovereignty, and the architecture that makes exit possible. The series follows the problem from its origins — the build-vs-b]]></description><link>https://nightthoughts.me/two-races-at-once-why-ai-is-both-a-security-emergency-and-a-competitive-one</link><guid isPermaLink="true">https://nightthoughts.me/two-races-at-once-why-ai-is-both-a-security-emergency-and-a-competitive-one</guid><category><![CDATA[finance]]></category><category><![CDATA[Cloud Computing]]></category><category><![CDATA[cyber security]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Fri, 02 Oct 2026 13:43:56 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/cf8d2a11-6092-4dc3-a352-9899799ede92.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is the second chapter of <strong>Built to Leave</strong>, a series on cloud dependency, sovereignty, and the architecture that makes exit possible. The series follows the problem from its origins — the build-vs-buy trade-off that predates AI — through the structural costs of cloud dependency, the work required to regain control, and the architectural shift that AI workloads are now forcing. Full table of contents: <a href="https://nightthoughts.hashnode.dev/build-to-leave">nightthoughts.hashnode.dev/build-to-leave</a></em></p>
<hr />
<p><a href="https://nightthoughts.hashnode.dev/the-old-trade-off-vendor-vs-in-house-before-ai-changed-the-rules">The previous piece</a> ended with two forces pulling financial institutions toward the same infrastructure, for very different reasons: one defensive, one competitive. This piece looks at what that means once both land on the same roadmap at the same time, based on what I've seen in cloud transformation work, kept at the general level rather than tied to any one employer.</p>
<h2>Security is not one bank's problem</h2>
<p>Security in a bank is never negotiable. A serious breach doesn't stay contained — banks are wired into each other, and a failure at one institution can ripple into all of them. That's why the sector's response is increasingly collective, not institution by institution. Project Glasswing, launched by Anthropic in April 2026, brings major banks and hyperscalers together — JPMorgan, AWS, Google, Microsoft, Palo Alto Networks — to jointly use AI to find and fix vulnerabilities before attackers do, and by June 2026 it had expanded to roughly 150 more partners across 15+ countries.</p>
<p>The urgency behind that effort just became a lot more concrete. On September 3, 2026, OpenAI released GPT-6 Astra — the first commercial AI model to reach OpenAI's own "Critical" threshold for cybersecurity capability. According to OpenAI's system card, Astra can find previously unknown security flaws and develop new ways to exploit them across well-protected systems, without a person guiding each step. In pre-release testing it discovered two zero-day vulnerabilities on its own. This is no longer hypothetical. As of this month, that capability exists commercially, under gated access. Fighting it requires AI-capable defense, alongside the discipline that comes with it: zero trust as a default, continuous scanning across binaries, source code, third-party libraries, and internal systems, and relentless tracking of vulnerabilities before they can be weaponized.</p>
<h2>The other pressure</h2>
<p>Security isn't the only force reshaping the roadmap. A second pressure is building in parallel, and it decides who's still standing in five years: competitiveness. There are two different things "AI" can mean here. One is tooling — copilots inside email, code, presentations — which makes existing work faster without changing what the business does. The other is AI as a business function: pricing, credit scoring, fraud detection — AI that changes the product itself. That second category is what actually moves the competitive needle.</p>
<p>Institutions that adopt it early are already pulling ahead on cost and performance; slower movers risk a cost base they can't compete with. Multiple surveys show banks now ranking cost reduction and productivity ahead of security as reasons to invest in AI. Even inside institutions that treat security as non-negotiable, the case for AI is increasingly built around staying competitive, not just staying safe.</p>
<h2>Where I've seen this actually play out</h2>
<p>The technology was rarely the hardest part. Two things running in parallel were.</p>
<p>The first is applying real security discipline against a technology landscape that, in most institutions I've seen, was never built with that discipline in mind. You can't zero-trust your way through hundreds of small, loosely documented applications overnight, especially against an adversary that can now automate vulnerability discovery at machine speed. Every one of those applications is its own scanning target, its own audit trail, its own zero-day exposure to track. Multiply a hard problem by a fragmented estate, and it becomes close to unmanageable regardless of budget.</p>
<p>The second is organizational: getting an institution to agree, in practice and not just in principle, on what gets funded first when security and AI competitiveness are both labeled the top priority. They compete for the same architects, the same budget cycle, the same production windows. That tension doesn't resolve through better slogans. It resolves by removing the reason it exists.</p>
<h2>The fix nobody wants to do first</h2>
<p>Most institutions aren't failing to prioritize security or competitiveness. They're trying to do both against an application landscape that should never have gotten this fragmented. Every small, tactically-built application — built to hit a deadline, wrapped around</p>
]]></content:encoded></item><item><title><![CDATA[The Old Trade-Off: Vendor vs. In-House, Before AI Changed the Rules]]></title><description><![CDATA[This is the first chapter of Built to Leave, a series on cloud dependency, sovereignty, and the architecture that makes exit possible. The series follows the problem from its origins — the build-vs-bu]]></description><link>https://nightthoughts.me/the-old-trade-off-vendor-vs-in-house-before-ai-changed-the-rules</link><guid isPermaLink="true">https://nightthoughts.me/the-old-trade-off-vendor-vs-in-house-before-ai-changed-the-rules</guid><category><![CDATA[AI]]></category><category><![CDATA[finance]]></category><category><![CDATA[Public Cloud]]></category><category><![CDATA[vendor]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Fri, 02 Oct 2026 13:37:43 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/2a6f1576-3f47-4503-a1fe-f38d0b0503ee.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is the first chapter of <strong>Built to Leave</strong>, a series on cloud dependency, sovereignty, and the architecture that makes exit possible. The series follows the problem from its origins — the build-vs-buy trade-off that predates AI — through the structural costs of cloud dependency, the work required to regain control, and the architectural shift that AI workloads are now forcing. Full table of contents: <a href="https://nightthoughts.hashnode.dev/build-to-leave">nightthoughts.hashnode.dev/build-to-leave</a></em></p>
<hr />
<p>Long before anyone in finance was talking about AI, institutions were already wrestling with one of the oldest questions in enterprise technology: build it yourself, or buy it from someone else. It's a decision every bank, insurer, and financial services firm has made repeatedly, across decades, for every new capability that came along — core banking systems, trading platforms, risk engines, data infrastructure. The technology changed. The underlying dilemma didn't.</p>
<p>This article isn't about AI yet. It's about understanding the trade-off as it existed before AI entered the picture — because you can't understand how AI is reshaping the game until you understand the board it's being played on.</p>
<h2>A brief chronology</h2>
<p>Financial institutions have swung between these two poles for a long time, and rarely by pure technical logic — often by cost pressure, regulation, or the scars of a previous decision gone wrong.</p>
<p>In earlier decades, many institutions built almost everything in-house, out of necessity more than preference — the vendor ecosystem for specialized financial technology simply didn't exist yet. Core systems were custom, maintained by large internal IT departments, because there was no market offering an alternative.</p>
<p>As the vendor landscape matured, the pendulum swung toward buying. Specialized providers emerged for almost every function — trading systems, compliance tooling, payment processing, data feeds — and outsourcing that work became the default, because building it yourself no longer made economic sense against a mature, specialized alternative.</p>
<p>More recently, the pendulum has been swinging back, at least partially. Institutions that outsourced heavily started reckoning with the risks of that dependency, and many began reinvesting in in-house engineering capability — not to replace vendors entirely, but to regain control over what they saw as strategically important.</p>
<p>Neither swing was permanent, and neither was universally correct. What's stayed constant is the underlying trade-off itself.</p>
<h2>The case for the vendor</h2>
<p>Buying from an external provider has always come with a clear pitch: <strong>you get expertise and speed without having to build the capability yourself.</strong></p>
<ul>
<li><p><strong>Cost predictability (at least on paper).</strong> You pay for a service rather than building and maintaining the underlying capability, converting a large fixed investment into a more variable one.</p>
</li>
<li><p><strong>Specialized expertise.</strong> Vendors often have deep, focused expertise in a narrow domain that would take years for an internal team to replicate.</p>
</li>
<li><p><strong>Faster time to market.</strong> A mature vendor product can often be deployed faster than a custom-built equivalent, because much of the engineering work is already done.</p>
</li>
<li><p><strong>Shared risk, in theory.</strong> SLAs and contracts create an expectation that the vendor shares responsibility for uptime, security, and support.</p>
</li>
</ul>
<p>But this pitch has always come with a quieter set of liabilities, ones that only become visible over time:</p>
<ul>
<li><p><strong>You don't control the roadmap.</strong> The vendor's priorities are not necessarily your priorities, and you may find yourself waiting on features, fixes, or changes that matter to you but not to their broader customer base.</p>
</li>
<li><p><strong>Dependency risk compounds.</strong> The longer you rely on a vendor, the more expensive and disruptive it becomes to leave — even if the relationship deteriorates.</p>
</li>
<li><p><strong>The vendor can fail you in ways outside your control.</strong> Contract terms can change unilaterally once switching costs are high enough. Vendors can be acquired, restructured, or go bankrupt, taking your roadmap down with them. A security breach or outage on their side becomes an incident on yours, regardless of whose code was vulnerable — regulators such as FINRA have specifically flagged the rise in cyberattacks and outages at third-party vendors as a risk that can ripple across large numbers of firms at once (<a href="https://www.finra.org/rules-guidance/guidance/reports/2025-finra-annual-regulatory-oversight-report/third-party-risk">FINRA</a>). And increasingly, geopolitical shifts — sanctions, export controls, data residency requirements — can sever access to a vendor's service almost overnight, regardless of what the contract says.</p>
</li>
<li><p><strong>You become the auditor, not the owner.</strong> Choosing a vendor doesn't remove your responsibility — regulators make that explicit in financial services. Outsourcing a function does not outsource the accountability for it: if a critical provider fails, the institution's own customers feel the consequences and its own regulator expects a response, including formal incident reporting where thresholds are met (<a href="https://www.calqrisk.com/resources/insights/outsourcing-and-third-party-risk-management-for-financial-firms">calQrisk</a>). The EU's Digital Operational Resilience Act (DORA), in force since January 2025, codifies exactly this: it gives supervisors the power to designate systemically important ICT and cloud providers as "critical third-party providers" subject to direct oversight, precisely because of the concentration risk created when many institutions lean on the same few providers (<a href="https://www.pifsinternational.org/cloud-adoption-in-the-financial-sector-and-concentration-risk/">PIFS International</a>; <a href="https://inquisitiveminds.bristows.com/post/102lqkb/aws-us-east-1-incident-regulators-concentrate-on-concentration-risk">Bristows</a>). Canada's OSFI guidance is just as direct, advising regulated institutions to consider strategies such as multi-cloud design specifically to mitigate cloud-provider concentration risk (<a href="https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/third-party-risk-management-guideline">OSFI</a>). That's the real substance behind "you become the auditor" — ongoing due diligence, SLA monitoring, security assessments, and resisting the temptation to let any single vendor become a single point of failure.</p>
</li>
</ul>
<h2>The case for in-house</h2>
<p>Building internally has always promised the opposite: <strong>control, at a price.</strong></p>
<ul>
<li><p><strong>Full ownership of the stack.</strong> No dependency on a third party's roadmap, pricing decisions, or continued existence.</p>
</li>
<li><p><strong>Institutional knowledge stays inside the company</strong>, compounding over time rather than walking out the door when a vendor contract ends.</p>
</li>
<li><p><strong>Tighter alignment with regulatory and compliance requirements</strong>, since governance sits entirely within the institution's own control rather than being negotiated through a vendor relationship.</p>
</li>
<li><p><strong>Faster internal feedback loops</strong>, at least once a team is mature — the people who build a system are also the ones maintaining and evolving it, without a vendor relationship as an intermediary.</p>
</li>
</ul>
<p>And the costs, while less visible than a vendor invoice, are just as real:</p>
<ul>
<li><p><strong>Fixed cost that doesn't scale down.</strong> Infrastructure and talent are ongoing commitments, whether the system is running at 20% capacity or 90%.</p>
</li>
<li><p><strong>Slower delivery, particularly in regulated environments.</strong> Every internal change typically passes through governance, audit trail requirements, and compliance review — overhead that a mature vendor product may have already absorbed into its design and certifications.</p>
</li>
<li><p><strong>Mission drift.</strong> For a financial institution, IT is not the core business. The more capability is brought in-house, the more the institution is, in effect, also becoming a software organization — competing for technical talent, carrying technical debt, and needing a level of engineering maturity that doesn't come naturally to a business built around finance, not software.</p>
</li>
<li><p><strong>Self-auditing is still auditing</strong> — <strong>and it's easier to under-invest in.</strong> Owning the stack doesn't remove the need for rigorous internal controls; it just moves that function inside the institution's own walls, where there's no external contract forcing the discipline. In-house risk tends to fail quietly, through neglect, while vendor risk tends to fail loudly, through breach disclosures and contract disputes — and that asymmetry in visibility often shapes the debate more than the actual underlying risk does.</p>
</li>
<li><p><strong>The operational headache no slide deck shows.</strong> Owning your IT end-to-end means owning problems that have nothing to do with writing good software. Running global operations means organizing a genuine follow-the-sun strategy — coverage across time zones, handovers that don't lose context, incident response that works at 3 a.m. somewhere no matter what. It means spreading knowledge across a distributed workforce rather than concentrating it in a single team, which is harder to do well than it sounds. And paradoxically, the more IT you own, the closer IT has to sit to the business — which means growing the very "bigger IT" organization that the mission-drift argument warns against in the first place.</p>
</li>
<li><p><strong>Data center provisioning is its own discipline.</strong> Standing up and running physical infrastructure — capacity planning, power, cooling, hardware refresh cycles, physical security — is a specialized competency in itself, largely unrelated to financial services. Institutions that own this end-to-end are committing real, ongoing effort to keep SLAs at the level the business requires, effort that has nothing to do with banking or finance and everything to do with running a data center well.</p>
</li>
<li><p><strong>Hosting costs that only go one direction.</strong> The cost of maintaining and upgrading data center infrastructure tends to climb over time, not fall. Construction costs per megawatt of IT load have risen sharply as demand for power-hungry, high-density facilities has surged, with electrical infrastructure now consuming close to half of total project budgets (<a href="https://terrapincg.com/news/average-cost-to-build-a-data-center-in-the-usa">Terrapin CG</a>; <a href="https://www.irecruit.co/insights/data-center-construction-cost-trends-2026">iRecruit</a>). Power itself has become the binding constraint: grid operators and energy researchers point to data center demand as a meaningful driver of rising electricity costs in several markets (<a href="https://fortune.com/2026/07/26/data-centers-electricity-costs-cheaper-7billion-buildout-ai-demand/">Fortune</a>). Running infrastructure at the resilience level financial regulators expect isn't cheap, and the trend line isn't moving in the direction that makes self-hosting easier over time.</p>
</li>
<li><p><strong>Resiliency requirements are escalating faster than most in-house setups can absorb.</strong> Geopolitical instability and worsening climate conditions — more frequent flooding, wildfires, earthquakes, and extreme weather events — are pushing infrastructure resilience requirements toward multi-region, geographically diversified setups almost as a baseline expectation now, not a best practice for the ambitious. This isn't a theoretical concern: one global analysis of nearly 9,000 data centers found that a meaningful share already sit in locations facing high or moderate physical climate risk, with several major hubs projected to see a large proportion of their facilities exposed by 2050 (<a href="https://www.datacenterdynamics.com/en/news/climate-threats-to-data-centers-set-to-surge-report/">DCD</a>; <a href="https://xdi.systems/news/global-data-centres-face-rising-climate-risks-xdi-report-warns-landmark-analysis-of-nearly-9000-sites-reveals-escalating-threat-to-digital-infrastructure">XDI</a>). A separate analysis of global data center markets put the share of capacity exposed to acute climate hazards — flooding, extreme wind, wildfire — at close to four in five facilities worldwide (<a href="https://www.cnbc.com/2026/06/18/data-center-climate-change-study.html">CNBC</a>). Meeting a rising resilience bar in-house means duplicating infrastructure across regions, at a cost and complexity that grows with every added region. At a certain point, this pressure alone pushes institutions to conclude that running physical infrastructure at that level of resiliency is better delegated to an organization for whom that <em>is</em> the core business — which is precisely the case public cloud providers make.</p>
</li>
</ul>
<h2>The speed trap: how both paths lead to the same mess</h2>
<p>There's a dimension to this trade-off that doesn't show up in cost comparisons or risk matrices, but that anyone who has actually shipped software in a bank will recognize immediately: speed pressure distorts both paths, and it tends to distort them toward the same outcome.</p>
<p>When the business sees an opportunity, the instinct is to move before competitors do. On the in-house side, that pressure produces what I'd call <strong>commando mode</strong> — or sniper mode: build exactly what's needed, nothing more, and move fast. No time for the full picture. No time to future-proof. Just enough to seize the opportunity, now, while it's there. And for day one, this works. It's often the <em>right</em> call — a mature, fully governed build process would have missed the window entirely.</p>
<p>The problem is what happens next. As soon as that commando-mode build reaches production and the business succeeds, the very success creates pressure the system was never built to absorb. It doesn't scale. Clients want more than the minimal version can deliver. The business struggles to understand why something that clearly works can't simply be extended — surely if it does <em>this</em>, it can do <em>that</em> too? And underneath, the parts that were skipped in the rush — proper security, authentication, resilience — are exactly the parts that are hardest and most expensive to retrofit once real usage and real data are flowing through the system.</p>
<p>The vendor path isn't immune to its own version of this trap, just from a different direction. Getting real value out of a vendor platform — configuring it properly, building the features specific to your business — requires expertise, and that expertise is often scarce and expensive. Institutions frequently end up dependent on a small number of specialists who understand the vendor's tool deeply, and just as often end up building a growing constellation of <strong>satellite applications</strong> around the vendor platform to bridge the gaps it doesn't cover natively. Ironically, by the time all of that is stitched together, time to market on the vendor path can end up <em>slower</em> than a commando-mode in-house build — the very speed advantage vendors are supposed to offer gets eaten by integration and customization overhead.</p>
<p>Here's the uncomfortable convergence: <strong>whichever path you take, the default outcome, without deliberate intervention, is the same</strong> — a sprawling landscape of small applications in production, built under time pressure, neither scalable nor properly secured, each one requiring real ongoing effort, spread across teams worldwide, just to keep it up and running.</p>
<p>This is where I think the real discipline lives — not in choosing build or buy correctly in the abstract, but in resisting the urge to let either path run on pure speed indefinitely. At some point, an institution has to deliberately slow down: organize what's been built, structure it properly, size the actual opportunity rather than the emergency that created it, and identify which of these applications are mature and important enough to deserve real investment — and which should be retired, consolidated, or never should have reached production in their current form in the first place.</p>
<h2>Neither side is "safe"</h2>
<p>What's striking, looking back across this history, is how often the debate gets framed as if one side is the "risky" choice and the other is the "safe" one. It never has been. Buying and building simply distribute risk differently — one concentrates it in a third-party relationship you have to actively manage, the other concentrates it inside your own walls, where it's your discipline alone keeping it in check.</p>
<p>The institutions that navigated this well weren't the ones that picked a permanent side. They were the ones that got specific: deciding deliberately which capabilities were core enough to own directly regardless of cost, and which were commodity enough that well-managed vendor risk was the better bet — and revisiting that decision as circumstances changed, rather than defaulting to "how we've always done it."</p>
<p>That balance held, imperfectly but functionally, for decades. But something is changing the terrain underneath it — and not gradually, the way most technology shifts do.</p>
<p>Part of the pressure is infrastructural: resiliency expectations are rising faster than most institutions can build for on their own, quietly pushing even committed in-house organizations toward providers whose entire business is running resilient infrastructure at global scale. Part of it is a discipline problem that predates any of this — the speed trap that fills production with small, fragile applications regardless of which path built them. And part of it is genuinely new, arriving from two directions at once.</p>
<p>One is defensive: the attacker on the other side of the wall is no longer just a person trying a few things and giving up when they fail. The other is competitive, and arguably the more urgent of the two: institutions that learn to use AI well are already starting to produce things faster, at lower cost, with better products, than those that don't. That's not a security question at all — it's a market-share question. An institution can have flawless security and still lose ground simply by being slower and more expensive than a competitor who figured out how to fold AI into their delivery cycle.</p>
<p>Both pressures point toward the same infrastructure — the data, compute, and platforms that make AI viable at scale — but for very different reasons. One is about not losing what you have. The other is about not falling behind. A few open questions to sit with before the next piece:</p>
<ul>
<li><p>If resiliency and infrastructure ownership are already pushing institutions toward the cloud on their own, what happens when both defending against an always-on attacker <em>and</em> keeping up with AI-native competitors make that pull even harder to resist?</p>
</li>
<li><p>Is the "slow down and get organized" discipline this article calls for even compatible with a competitive environment where the institutions moving fastest with AI are also the ones pulling ahead?</p>
</li>
<li><p>And if the answer increasingly runs through public cloud — whose infrastructure is it, really, and what does that mean for an institution's independence when both its security and its competitiveness depend on it?</p>
</li>
</ul>
<p><em>What happens to this trade-off when infrastructure itself becomes too demanding to own alone, discipline becomes harder to enforce under competitive pressure, the attacker never sleeps, and standing still means falling behind? That's where we go next.</em></p>
<hr />
<h3>References</h3>
<ul>
<li><p>PIFS International, <a href="https://www.pifsinternational.org/cloud-adoption-in-the-financial-sector-and-concentration-risk/">"Cloud Adoption in the Financial Sector and Concentration Risk"</a></p>
</li>
<li><p>Office of the Superintendent of Financial Institutions (OSFI), <a href="https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/third-party-risk-management-guideline">"Third-Party Risk Management Guideline"</a></p>
</li>
<li><p>Bristows, <a href="https://inquisitiveminds.bristows.com/post/102lqkb/aws-us-east-1-incident-regulators-concentrate-on-concentration-risk">"AWS US-EAST-1 incident: regulators concentrate on concentration risk"</a></p>
</li>
<li><p>calQrisk, <a href="https://www.calqrisk.com/resources/insights/outsourcing-and-third-party-risk-management-for-financial-firms">"Outsourcing and Third-Party Risk Management for Financial Firms"</a></p>
</li>
<li><p>FINRA, <a href="https://www.finra.org/rules-guidance/guidance/reports/2025-finra-annual-regulatory-oversight-report/third-party-risk">"Third-Party Risk Landscape," 2025 Annual Regulatory Oversight Report</a></p>
</li>
<li><p>Data Center Dynamics, <a href="https://www.datacenterdynamics.com/en/news/climate-threats-to-data-centers-set-to-surge-report/">"Climate threats to data centers set to surge"</a></p>
</li>
<li><p>XDI, <a href="https://xdi.systems/news/global-data-centres-face-rising-climate-risks-xdi-report-warns-landmark-analysis-of-nearly-9000-sites-reveals-escalating-threat-to-digital-infrastructure/">"Global data centres face rising climate risks"</a></p>
</li>
<li><p>CNBC, <a href="https://www.cnbc.com/2026/06/18/data-center-climate-change-study.html">"Data center climate change study"</a></p>
</li>
<li><p>Terrapin CG, <a href="https://terrapincg.com/news/average-cost-to-build-a-data-center-in-the-usa">"Average Cost to Build a Data Center in the USA"</a></p>
</li>
<li><p>iRecruit, <a href="https://www.irecruit.co/insights/data-center-construction-cost-trends-2026">"Data Center Construction Cost Trends 2026"</a></p>
</li>
<li><p>Fortune, <a href="https://fortune.com/2026/07/26/data-centers-electricity-costs-cheaper-7billion-buildout-ai-demand/">"Data centers and electricity costs"</a></p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[The Secret Digital Workers Running the Internet: A Guide to Non-Human Identities]]></title><description><![CDATA[Imagine walking into a modern digital hotel. You use a plastic key card to unlock your room door. That key card represents your human identity—it proves who you are and gives you permission to enter. ]]></description><link>https://nightthoughts.me/the-secret-digital-workers-running-the-internet-a-guide-to-non-human-identities</link><guid isPermaLink="true">https://nightthoughts.me/the-secret-digital-workers-running-the-internet-a-guide-to-non-human-identities</guid><category><![CDATA[AI]]></category><category><![CDATA[ai agents]]></category><category><![CDATA[cybersecurity]]></category><category><![CDATA[Security]]></category><category><![CDATA[Cloud]]></category><category><![CDATA[DevSecOps]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Tue, 29 Sep 2026 22:14:41 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/705f741d-1fea-4644-86cb-c7fc33bd9896.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine walking into a modern digital hotel. You use a plastic key card to unlock your room door. That key card represents your <strong>human identity</strong>—it proves who you are and gives you permission to enter. </p>
<p>However, behind the scenes, hundreds of automated systems are moving without human intervention: luggage delivery bots unlock elevator doors, security cameras check authority passes, and automated payment kiosks talk to bank servers. </p>
<p>To open those doors, these machines need their own keys.</p>
<p>In the digital world, humans are no longer the only ones logging in. Computer programs, automated scripts, software tools, and AI agents log in millions of times per second. These machine passes are called <strong>Non-Human Identities (NHIs)</strong>.</p>
<p>In fact, in modern enterprise environments, non-human identities outnumber human identities by a ratio ranging between <strong>45:1 and 140:1</strong> <a href="#ref-1">1</a>.</p>
<hr />
<h2>What Exactly is a Non-Human Identity (NHI)?</h2>
<p>At its simplest, a <strong>Non-Human Identity</strong> is a set of digital credentials—such as an API key, a service account, or a digital token—that allows one software program to speak to another program without a human having to type in a password <a href="#ref-2">2</a>.</p>
<ul>
<li><strong>An API Key:</strong> A special digital password that lets your weather phone app fetch live updates from a weather server.</li>
<li><strong>A Service Account:</strong> A background login created so a cloud service can automatically back up your files at midnight.</li>
<li><strong>A Digital Certificate:</strong> An encrypted pass that proves a website or device is legitimate and safe to connect with.</li>
</ul>
<p>Without NHIs, the modern internet would grind to a halt because humans would have to manually approve every single background data exchange.</p>
<hr />
<h2>The 4 Places Where Unmanaged Machine Keys Hide</h2>
<p>Because these machine keys are created automatically, they are easy to forget. Security teams face four major challenges when trying to manage them:</p>
<ol>
<li><strong>Cloud-Native Growth:</strong> As companies move to the cloud, billions of tiny software connections are made across multiple servers <a href="#ref-2">2</a>.</li>
<li><strong>Software Assembly Lines (CI/CD Pipelines):</strong> Developers use automated tools to build and test code continuously. At every step of this automated assembly line, software tools leave behind digital access keys <a href="#ref-1">1</a>.</li>
<li><strong>Smart Devices &amp; Forgotten Systems (IoT):</strong> Connected hardware, old dormant accounts, and expired digital certificates sit silently in networks, still holding active master access <a href="#ref-2">2</a>.</li>
<li><strong>Supply Chain Connections:</strong> Modern apps rely on third-party vendor tools. If an app trusts a vendor's API, it creates <strong>inherited trust</strong>. If that vendor gets hacked, the attacker can walk right into your system through that trusted connection <a href="#ref-1">1</a>.</li>
</ol>
<hr />
<h2>The Emerging Threat: Logging In with "Borrowed Trust"</h2>
<p>Security leaders have noticed a major shift: <strong>80% of security leaders rank AI and machine-related identity risks as their top concern</strong> <a href="#ref-3">3</a>.</p>
<p>Why? Because tricking a human into giving up a password takes effort, and humans often have two-factor authentication (like a text message code). </p>
<p>Machine keys, however, rarely have two-factor checks. If a hacker finds a forgotten API key accidentally published in code or stored on a server, they don't need to break in. They simply <strong>"log in" using borrowed trust</strong>—the system assumes the attacker is just another friendly internal software tool <a href="#ref-2">2</a>.</p>
<hr />
<h2>The Agentic AI Escalation</h2>
<p>The risk becomes even bigger with the rise of <strong>Agentic AI</strong>. </p>
<p>Traditional software only follows rigid, step-by-step instructions. <strong>Agentic AI</strong>, by contrast, makes autonomous decisions to complete complex goals <a href="#ref-2">2</a>:</p>
<ul>
<li>It can execute financial transactions.</li>
<li>It can manage supply chain inventory and place orders automatically.</li>
<li>It can interact directly with industrial control systems.</li>
</ul>
<h3>The Danger: Context Loss in Multi-Agent Chains</h3>
<p>When AI Agent A hires AI Agent B, which then triggers AI Agent C to perform a financial trade, software systems can experience <strong>context loss</strong>. They lose track of who originally authorized the command, making it easy for mistakes or malicious instructions to slip through unnoticed <a href="#ref-2">2</a>.</p>
<hr />
<h2>The WEF Safety Framework for Machine Identities</h2>
<p>To help organizations protect themselves, cybersecurity experts at the <strong>World Economic Forum (WEF)</strong> and its Global Future Councils designed a five-step governance framework specifically for non-human identities and agentic AI <a href="#ref-2">2</a>:</p>
<blockquote>
<p><strong>1. Universal Discovery</strong> ➔ <strong>2. Eliminate Static Keys</strong> ➔ <strong>3. Extend Zero Trust</strong> ➔ <strong>4. Behavior Anomaly Detection</strong> ➔ <strong>5. Identity Delegation Tracing</strong></p>
</blockquote>
<h3>1. Universal Discovery &amp; Ownership</h3>
<p>You cannot protect what you cannot see. The WEF framework emphasizes that organizations must inventory every single API key, bot, and AI agent, and assign a <strong>responsible human owner</strong> to every machine identity <a href="#ref-2">2</a>.</p>
<h3>2. Eliminate Long-Lived Secrets</h3>
<p>Static passwords that last for years are dangerous. Organizations must phase out static keys and replace them with <strong>ephemeral (short-lived) tokens</strong> that expire automatically after a few minutes or hours <a href="#ref-2">2</a>.</p>
<h3>3. Extend "Zero Trust" to Machines</h3>
<p>Never assume a program is safe just because it is inside your network <a href="#ref-2">2</a>.</p>
<ul>
<li><strong>Continuous Authorization:</strong> Continuously re-verify machine identities.</li>
<li><strong>Strict Least-Privilege Scoping:</strong> Give a machine key access <em>only</em> to the specific folder or task it needs—nothing more.</li>
<li><strong>Post-Quantum Cryptography:</strong> Upgrade security standards so future quantum computers won't be able to crack machine keys <a href="#ref-2">2</a>.</li>
</ul>
<h3>4. Behavior-Based Anomaly Detection</h3>
<p>Instead of just checking if a key is valid when it logs in, systems must watch <strong>how the machine behaves</strong>. If a bot that normally checks weather updates suddenly attempts to download a customer database at 3 AM, the system must immediately flag and block it <a href="#ref-2">2</a>.</p>
<h3>5. Identity Delegation Tracing</h3>
<p>As AI agents pass commands down a chain of other tools, temporary safety passes must be attached at every step. This creates a full <strong>audit trail</strong> so humans can always review exactly why an AI agent made a specific decision <a href="#ref-2">2</a>.</p>
<hr />
<h2>Summary</h2>
<p>As AI shifts from simple assistants to autonomous agents that take action on our behalf, protecting machine keys is no longer just a technical detail—it is the foundation of digital safety. Frameworks like the World Economic Forum's give us a roadmap to manage these digital workers safely as our systems grow.</p>
<hr />
<h2>References</h2>
<p><a id="ref-1"></a>
<strong>[1] Enterprise NHI Statistics (45:1 to 140:1 Ratio):</strong> <a href="https://cloudsecurityalliance.org">Cloud Security Alliance (CSA)</a> &amp; <a href="https://entro.security">Entro Security Research</a>. Data shows non-human identities outnumber human users by 45:1 in average enterprise setups and over 140:1 in cloud-native environments.</p>
<p><a id="ref-2"></a>
<strong>[2] World Economic Forum (WEF) NHI Framework:</strong> <a href="https://www.weforum.org/centre-for-cybersecurity">World Economic Forum Centre for Cybersecurity</a>. Research and governance guidelines published by the Global Future Councils, titled <em>"Governing Non-Human Identities &amp; Agentic AI."</em></p>
<p><a id="ref-3"></a>
<strong>[3] CISO Threat Landscape Surveys:</strong> <a href="https://www.gartner.com/en/information-technology/role/ciso-cybersecurity-leaders">Gartner Cybersecurity &amp; CISO Insights</a>. Global Chief Information Security Officer (CISO) Industry Reports on Identity &amp; Access Management (IAM), showing ~80% of security executives rank non-human identity exposure and autonomous AI access as top priorities.</p>
]]></content:encoded></item><item><title><![CDATA[Traditional Vector RAG vs. PageIndex: What's the Difference?]]></title><description><![CDATA[Imagine you need to answer a specific question using a 200-page financial report.
How would you do it? You would open the document, look at the Table of Contents, find the right chapter, turn to those]]></description><link>https://nightthoughts.me/traditional-vector-rag-vs-pageindex-what-s-the-difference</link><guid isPermaLink="true">https://nightthoughts.me/traditional-vector-rag-vs-pageindex-what-s-the-difference</guid><category><![CDATA[AI]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[RAG ]]></category><category><![CDATA[llm]]></category><category><![CDATA[Open Source]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Tue, 29 Sep 2026 21:26:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/5765d933-734e-47ed-ab79-0ca0c95e2883.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine you need to answer a specific question using a 200-page financial report.</p>
<p>How would you do it? You would open the document, look at the Table of Contents, find the right chapter, turn to those pages, and read the entire section to understand the full context.</p>
<p>Most standard AI search systems do not read documents this way. Instead, they use a process called <strong>Retrieval-Augmented Generation (RAG)</strong>, which cuts documents into tiny, disconnected pieces before searching them.</p>
<p>An open-source framework called <strong>PageIndex</strong> changes this approach. Here is a clear breakdown of how traditional Vector RAG works, how PageIndex offers a vectorless alternative, and where each approach shines.</p>
<hr />
<h2>What is PageIndex and Who Created It?</h2>
<p><strong>PageIndex</strong> is an open-source framework designed to give AI models a better way to search and analyze long documents without using vector databases or fixed text chunking. It is known as a <strong>vectorless, reasoning-based RAG engine</strong>.</p>
<p>PageIndex was created by the team at <strong>Vectify AI</strong> (led by researchers including Mingtian Zhang and Yu Tang) and released on GitHub.</p>
<p>Instead of converting text into abstract mathematical numbers (vectors), PageIndex converts documents into a <strong>hierarchical tree structure</strong>—essentially a detailed, machine-readable Table of Contents complete with page boundaries and section summaries.</p>
<hr />
<h2>Traditional Vector RAG vs. PageIndex: Step-by-Step</h2>
<p>To understand how PageIndex changes document retrieval, compare how both systems process a complex document.</p>
<table>
<thead>
<tr>
<th>Process Step</th>
<th>Traditional Vector RAG</th>
<th>PageIndex (Vectorless RAG)</th>
</tr>
</thead>
<tbody><tr>
<td><strong>1. Document Ingestion</strong></td>
<td>Cuts documents into small, fixed-size <strong>chunks</strong> (e.g., 512 words).</td>
<td>Extracts the document's natural <strong>structure</strong> (chapters, headers, subsections).</td>
</tr>
<tr>
<td><strong>2. Information Indexing</strong></td>
<td>Converts text chunks into mathematical numbers stored in a <strong>Vector DB</strong>.</td>
<td>Builds a hierarchical <strong>Tree Index</strong> with node summaries and exact page/line coordinates.</td>
</tr>
<tr>
<td><strong>3. Query Retrieval</strong></td>
<td>Compares query vectors against chunk vectors using mathematical distance.</td>
<td>An AI model <strong>reasons</strong> through the tree branch-by-branch to locate the target section.</td>
</tr>
<tr>
<td><strong>4. Final Context Delivered</strong></td>
<td>Pulls 1–2 <strong>isolated text snippets</strong> without surrounding context.</td>
<td>Pulls <strong>complete contiguous sections/pages</strong> with exact page citations.</td>
</tr>
</tbody></table>
<hr />
<h2>Key Benefits of PageIndex</h2>
<ul>
<li><p><strong>Preserves Full Context:</strong> Traditional RAG often cuts a sentence or table in half across chunk boundaries. PageIndex retrieves contiguous physical pages, ensuring the AI model sees the complete thought.</p>
</li>
<li><p><strong>High Traceability:</strong> Every answer points back to an exact section and page range (e.g., "Pages 12–14"), making it straightforward to double-check source facts.</p>
</li>
<li><p><strong>No Vector Database Infrastructure:</strong> You do not need to set up, maintain, or manage external vector databases like Pinecone or ChromaDB.</p>
</li>
<li><p><strong>Higher Accuracy on Complex Docs:</strong> On structured reports, reasoning through section hierarchies yields significantly higher answer accuracy than matching isolated keywords or phrases.</p>
</li>
</ul>
<hr />
<h2>Where PageIndex Isn't Perfect: The Limitations</h2>
<p>While PageIndex fixes major flaws in traditional chunk-based retrieval, it is not a solution for every problem. It comes with clear trade-offs:</p>
<ul>
<li><p><strong>Higher Latency (Slower Retrieval):</strong> Because PageIndex relies on an LLM to read through tree nodes and navigate branches step-by-step, queries take longer than fast vector math lookups.</p>
</li>
<li><p><strong>Higher API Token Costs:</strong> Traversing node summaries with an LLM sends more tokens back and forth, which increases the execution cost per query.</p>
</li>
<li><p><strong>Requires Structured Documents:</strong> PageIndex relies on documents having a logical layout (headers, titles, clear sections). It struggles with completely unstructured text, poor OCR scans, or unorganized text files.</p>
</li>
<li><p><strong>Not Ideal for Massive Multi-Document Pools:</strong> If you need to search across 100,000 separate 1-page customer support emails simultaneously, traditional vector search remains much faster and more scalable.</p>
</li>
</ul>
<hr />
<h2>Choosing the Right Tool for Your Project</h2>
<p>Selecting between Traditional Vector RAG and PageIndex depends on your data structure and query requirements.</p>
<p>Use <strong>PageIndex</strong> when working with long, structured documents—such as financial 10-K filings, legal contracts, technical manuals, or research papers—where accuracy, full context, and exact page citations are required.</p>
<p>Use <strong>Traditional Vector RAG</strong> when you need sub-second search speeds across thousands of short, unstructured files, or where low latency and lower token costs outweigh the need for deep contextual reasoning.</p>
<p>To explore the codebase and run a quickstart implementation, visit the official repository at <code>VectifyAI/PageIndex</code> on GitHub.</p>
]]></content:encoded></item><item><title><![CDATA[How to Build an Intelligent Router for NVIDIA and Ollama with LiteLLM, JEV, and Memori on a Mac Mini (Part 2)]]></title><description><![CDATA[Part 1 covered the foundation: an NVIDIA API key, a local proxy for NVIDIA NIM, and OpenClaw configured with its native NVIDIA provider. That setup gave the Mac mini access to frontier NVIDIA models, ]]></description><link>https://nightthoughts.me/how-to-build-an-intelligent-router-for-nvidia-and-ollama-with-litellm-jev-and-memori-on-a-mac-mini-part-2</link><guid isPermaLink="true">https://nightthoughts.me/how-to-build-an-intelligent-router-for-nvidia-and-ollama-with-litellm-jev-and-memori-on-a-mac-mini-part-2</guid><category><![CDATA[llm]]></category><category><![CDATA[NVIDIA]]></category><category><![CDATA[AI]]></category><category><![CDATA[local-models]]></category><category><![CDATA[litellm]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Mon, 28 Sep 2026 22:57:40 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/e48ef2bc-e51e-4032-b1e3-6db06ae7ec95.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Part 1 covered the foundation: an NVIDIA API key, a local proxy for NVIDIA NIM, and OpenClaw configured with its native NVIDIA provider. That setup gave the Mac mini access to frontier NVIDIA models, but it kept two separate paths. Local Ollama models lived on <code>localhost:11434</code>. NVIDIA models went through the proxy on <code>localhost:5002</code>. Clients had to know which backend to call.</p>
<p>Part 2 removes that choice. A single intelligent router sits in front of both Ollama and NVIDIA NIM, decides which model serves each request, and adds persistent memory across sessions.</p>
<p>If the foundation from Part 1 is not in place, it is covered here: <a href="https://nightthoughts.hashnode.dev/how-to-set-up-nvidia-api-keys-and-use-them-with-ollama-and-openclaw-on-a-mac-mini-part-1">How to Set Up NVIDIA API Keys and Use Them with Ollama and OpenClaw on a Mac Mini (Part 1)</a>.</p>
<h2>A Note on the Proxy from Part 1</h2>
<p>Part 1 used <code>ollama-proxy</code> to bridge Ollama and NVIDIA NIM. That bridge is no longer needed. LiteLLM handles the same job natively, along with provider abstraction, fallbacks, retries, and routing logic. The proxy becomes redundant.</p>
<p>Retire it once clients point to LiteLLM:</p>
<pre><code class="language-shell">launchctl unload ~/Library/LaunchAgents/com.ollama.proxy.plist
</code></pre>
<p>If any tools still call port <code>5002</code> with <code>nvidia-nim:</code> prefixed models, keep the proxy running until those clients are updated. Then remove it. Running both long-term adds two configs to maintain and two places for the API key.</p>
<p>Part 1 was a valid stepping stone. Part 2 is the upgrade.</p>
<h2>The Stack at a Glance</h2>
<p>Three components work together to route between local Ollama models and NVIDIA NIM:</p>
<table>
<thead>
<tr>
<th>Component</th>
<th>Role</th>
</tr>
</thead>
<tbody><tr>
<td><strong>LiteLLM</strong></td>
<td>OpenAI-compatible proxy that routes to 100+ providers, including Ollama and NVIDIA NIM. Replaces <code>ollama-proxy</code> from Part 1.</td>
</tr>
<tr>
<td><strong>JEV</strong></td>
<td>TypeSafe's System One model. A lightweight decision layer that picks which model serves a request.</td>
</tr>
<tr>
<td><strong>Memori</strong></td>
<td>SQL-native memory engine that gives the router persistent, queryable memory across sessions.</td>
</tr>
</tbody></table>
<p>Each one is replaceable. LiteLLM handles transport and provider abstraction. JEV handles routing intelligence. Memori handles context and recall.</p>
<h2>Architecture</h2>
<pre><code class="language-html">flowchart TB
    A[Clients] --&gt; D[LiteLLM Router]
    D --&gt; E[JEV decision]
    E --&gt; F[Model pool]
    F --&gt; H[Local Ollama]
    F --&gt; I[NVIDIA NIM]
    D --&gt; G[Memori]
    G --&gt; D
</code></pre>
<p>The architecture consists of three layers:</p>
<ol>
<li><p><strong>Clients</strong> send requests to the LiteLLM router on port <code>4000</code>. Clients include OpenClaw, Ollama clients, and any OpenAI-compatible tool.</p>
</li>
<li><p><strong>The Router</strong> is powered by LiteLLM. It uses a pre-call hook to build a request summary. The JEV decision layer picks the cheapest eligible tier. The model pool routes the request to either local Ollama or NVIDIA NIM.</p>
</li>
<li><p><strong>Memory</strong> is provided by Memori. It connects to the pre-call hook. It recalls relevant memories before the call and stores the conversation after the call in a local SQLite database.</p>
</li>
</ol>
<p>The router listens on <code>http://localhost:4000</code>. Local Ollama runs on <code>localhost:11434</code>. NVIDIA NIM is reached at <code>https://integrate.api.nvidia.com/v1</code>.</p>
<h3>How a Request Flows</h3>
<ol>
<li><p><strong>Client sends a request</strong> to the LiteLLM proxy with the model set to <code>jev-router</code>.</p>
</li>
<li><p><strong>The pre-call hook</strong> intercepts the request and builds a minimized summary: recent message roles, truncated text, and signals for images or tool use.</p>
</li>
<li><p><strong>The decision layer</strong> picks which tier should serve the request. It returns a tier choice with confidence and probabilities.</p>
</li>
<li><p><strong>LiteLLM routes</strong> the request to the selected model in the pool. Local models go to Ollama. NVIDIA models go to NVIDIA NIM.</p>
</li>
<li><p><strong>Memori</strong> injects relevant memories before the call and stores the conversation afterward, using a standard SQL database as the backing store.</p>
</li>
</ol>
<h2>Why This Architecture</h2>
<h3>Single Endpoint</h3>
<p>Clients no longer need to know which backend to use. They call <code>http://localhost:4000/v1/chat/completions</code> with <code>model: "jev-router"</code>. LiteLLM and the decision layer handle the rest, routing to either Ollama or NVIDIA NIM.</p>
<h3>One Component Instead of Two</h3>
<p>Part 1 required a separate proxy for NVIDIA and direct calls to Ollama for local models. LiteLLM absorbs both. The NVIDIA API key lives in one place. The Ollama endpoint lives in one place. Clients see one URL.</p>
<h3>Intelligent Routing</h3>
<p>Prefix-based routing is simple but rigid. <code>nvidia-nim:</code> always goes to NVIDIA. <code>llama3.2</code> always goes to Ollama. A decision layer adds judgment. It looks at the request and picks the cheapest tier whose models can fully answer it.</p>
<h3>Persistent Memory</h3>
<p>Memori gives the router a memory that survives restarts. It works with any LLM framework and uses standard SQL databases instead of vector stores, which cuts memory costs by an estimated 80-90%.</p>
<h2>Security Principles</h2>
<p>Before installing anything, three decisions shape the security posture of the entire stack.</p>
<p><strong>Pin LiteLLM to a safe version.</strong> LiteLLM had critical vulnerabilities in 2026, including a supply chain incident in March and several CVEs patched in v1.83.0 and later. The version floor for any proxy is <strong>&gt;= 1.83.7</strong>. Version 1.83.7 fixes CVE-2026-42271, an arbitrary command execution vulnerability. Install with an exact pinned version and verify.</p>
<p><strong>Generate a strong master key.</strong> The LiteLLM master key is the proxy's admin credential. Without it, anyone who can reach port <code>4000</code> can use every model in the pool. The key must start with <code>sk-</code> and be generated from at least 32 random bytes. Never use the default <code>sk-1234</code> placeholder. Exposed gateways are a known problem.</p>
<p><strong>Bind to localhost only.</strong> LiteLLM binds to <code>0.0.0.0</code> by default, which exposes the proxy to the local network. Add <code>--host 127.0.0.1</code> to every run command and launchd plist.</p>
<p><strong>Use virtual keys for clients.</strong> The master key is the admin credential. It must never be handed to a consumer. Clients should use scoped virtual keys, so access can be revoked without rotating the master key.</p>
<p><strong>Disable environment credential login for the Admin UI.</strong> If the Admin UI is not needed, do not expose it. If it is needed, set <code>disable_env_credential_login: true</code> in <code>config.yaml</code> after creating a proper admin account.</p>
<p>These principles are applied step-by-step in the sections that follow.</p>
<h2>Before You Start: Verify NVIDIA API Key</h2>
<p>The <code>NVIDIA_NIM_API_KEYS</code> variable is set in Part 1. If it is already in your shell profile, skip to the next section. If not, or if you are starting fresh, verify it now.</p>
<p>Check the current shell:</p>
<pre><code class="language-shell">echo "NVIDIA_NIM_API_KEYS=${NVIDIA_NIM_API_KEYS:-unset}"
</code></pre>
<p>Expected output if set:</p>
<pre><code class="language-text">NVIDIA_NIM_API_KEYS=nvapi-xxxxxxxxxxxxxxxxxxxx
</code></pre>
<p>If it prints <code>unset</code>, add it to your shell profile. For Zsh:</p>
<pre><code class="language-shell">echo 'export NVIDIA_NIM_API_KEYS="nvapi-xxxxxxxxxxxxxxxxxxxx"' &gt;&gt; ~/.zshrc
source ~/.zshrc
</code></pre>
<p>For Bash:</p>
<pre><code class="language-shell">echo 'export NVIDIA_NIM_API_KEYS="nvapi-xxxxxxxxxxxxxxxxxxxx"' &gt;&gt; ~/.bash_profile
source ~/.bash_profile
</code></pre>
<p>Replace the placeholder with the actual key from <code>build.nvidia.com</code>. The proxy and LiteLLM both read this variable.</p>
<h2>Installing the Stack</h2>
<h3>Prerequisites</h3>
<p>The Mac mini already has Python 3.11 or later from Part 1. Verify:</p>
<pre><code class="language-shell">python3 --version
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">Python 3.11.x
</code></pre>
<h3>Install LiteLLM</h3>
<p>Install the pinned version with proxy support:</p>
<pre><code class="language-shell">pip3 install 'litellm[proxy]==1.83.7'
</code></pre>
<p>Verify:</p>
<pre><code class="language-shell">litellm --version
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">litellm, version 1.83.7
</code></pre>
<h3>Install Memori</h3>
<pre><code class="language-shell">pip3 install memorisdk
</code></pre>
<p>Create the memory directory:</p>
<pre><code class="language-shell">mkdir -p ~/ai-router/memory
</code></pre>
<h3>Install the JEV Router</h3>
<pre><code class="language-shell">pip3 install jev-router
</code></pre>
<p>Alternatively, clone the repository:</p>
<pre><code class="language-shell">git clone https://github.com/prismhq/jev-router.git ~/jev-router
cd ~/jev-router
pip3 install -e .
</code></pre>
<p>Verify the hook is available:</p>
<pre><code class="language-shell">python3 -c "import jev_router; print(jev_router.__version__)"
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">0.1.x
</code></pre>
<h3>Create the Router Directory</h3>
<pre><code class="language-shell">mkdir -p ~/ai-router/{config,memory,logs}
</code></pre>
<h2>Base LiteLLM Configuration</h2>
<p>Create <code>~/ai-router/config.yaml</code>. This file defines the model pool and the master key.</p>
<pre><code class="language-yaml">model_list:
  - model_name: local-llama
    litellm_params:
      model: ollama_chat/llama3.2
      api_base: http://localhost:11434

  - model_name: local-mistral
    litellm_params:
      model: ollama_chat/mistral:7b
      api_base: http://localhost:11434

  - model_name: nvidia-deepseek
    litellm_params:
      model: nvidia_nim/deepseek-ai/deepseek-v4.1-flash
      api_key: os.environ/NVIDIA_NIM_API_KEYS

  - model_name: nvidia-glm
    litellm_params:
      model: nvidia_nim/z-ai/glm-5-3
      api_key: os.environ/NVIDIA_NIM_API_KEYS

litellm_settings:
  callbacks: jev_router.hook:router_hook
  drop_params: true

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  disable_env_credential_login: true
</code></pre>
<p>The <code>callbacks</code> line registers the JEV router's pre-call hook. The <code>master_key</code> line enforces authentication. The <code>disable_env_credential_login</code> line prevents the Admin UI from accepting the default environment-based login.</p>
<p>Export the master key:</p>
<pre><code class="language-shell">echo 'export LITELLM_MASTER_KEY="sk-$(openssl rand -hex 32)"' &gt;&gt; ~/.zshrc
source ~/.zshrc
</code></pre>
<h3>Create a Virtual Key for OpenClaw</h3>
<p>Start the proxy manually for key generation:</p>
<pre><code class="language-shell">cd ~/ai-router
litellm --config config.yaml --host 127.0.0.1 --port 4000
</code></pre>
<p>In another terminal, generate a virtual key:</p>
<pre><code class="language-shell">curl -X POST http://localhost:4000/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "models": ["local-llama", "local-mistral", "nvidia-deepseek", "nvidia-glm"],
    "metadata": {"client": "openclaw"}
  }'
</code></pre>
<p>The response contains the virtual key. Store it with the client. Generate separate keys for other tools.</p>
<p>Stop the proxy with <code>Ctrl+C</code> once the key is saved.</p>
<h2>Router Configuration</h2>
<p>The JEV router ships with two decision engines: <code>JevDecider</code> and <code>RulesDecider</code>. The choice is automatic. If <code>TYPESAFE_API_KEY</code> is set, <code>JevDecider</code> runs and JEV makes the routing call. If the key is not set, <code>RulesDecider</code> runs instead. It picks the cheapest eligible model from the candidate pool, filtered by capability flags like vision, output length, and tool support.</p>
<p>For a local-first setup, the default is <code>RulesDecider</code>. No key, no third-party call, no data leaving the Mac mini.</p>
<h3>Ensuring Privacy: Unsetting <code>TYPESAFE_API_KEY</code></h3>
<p><code>RulesDecider</code> runs only when <code>TYPESAFE_API_KEY</code> is not present. If the variable is set anywhere — shell profile, <code>.env</code> file, or launchd plist — the router will use <code>JevDecider</code> instead and send minimized summaries to TypeSafe.</p>
<p>Verify the variable is not set:</p>
<pre><code class="language-shell">echo "TYPESAFE_API_KEY=${TYPESAFE_API_KEY:-unset}"
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">TYPESAFE_API_KEY=unset
</code></pre>
<p>If it prints a value, remove it from your shell profile:</p>
<pre><code class="language-shell">sed -i '' '/TYPESAFE_API_KEY/d' ~/.zshrc
source ~/.zshrc
</code></pre>
<p>Check any <code>.env</code> files:</p>
<pre><code class="language-shell">grep -R "TYPESAFE_API_KEY" ~/jev-router 2&gt;/dev/null
</code></pre>
<p>Check the launchd plist:</p>
<pre><code class="language-shell">grep -A 10 "EnvironmentVariables" ~/Library/LaunchAgents/com.litellm.router.plist
</code></pre>
<p>If the variable is listed, remove it. Reload the service if it was running.</p>
<h3>Configure the Routing Policy</h3>
<p>Create <code>~/ai-router/router.yaml</code>. This file defines the candidates, their capabilities, prices, and the fallback.</p>
<pre><code class="language-yaml">candidates:
  - name: local-llama
    description: Fast local model for simple tasks.
    price: 0.0
    capabilities:
      vision: false
      tools: true
      max_output: 4096

  - name: local-mistral
    description: Balanced local model for everyday coding.
    price: 0.0
    capabilities:
      vision: false
      tools: true
      max_output: 8192

  - name: nvidia-deepseek
    description: NVIDIA NIM model for complex reasoning and long context.
    price: 0.0001
    capabilities:
      vision: true
      tools: true
      max_output: 131072

  - name: nvidia-glm
    description: NVIDIA NIM model for multimodal and agentic tasks.
    price: 0.0001
    capabilities:
      vision: true
      tools: true
      max_output: 131072

fallback: local-mistral
</code></pre>
<p>Prices are relative. <code>RulesDecider</code> picks the cheapest model that meets the request's capability needs. Local models have price <code>0.0</code>, so they win for any request they can handle. NVIDIA models only get selected when the request needs vision, tools, or a context window larger than the local models support.</p>
<h3>Switching to Another Routing Path</h3>
<p>The endpoint stays <code>http://localhost:4000/v1/chat/completions</code>. The model name stays <code>jev-router</code>. Only the decision layer changes.</p>
<p><strong>Path 2: Local Classifier</strong> replaces <code>RulesDecider</code> with a small Ollama model that reads the prompt and returns a tier.</p>
<pre><code class="language-yaml">model_list:
  - model_name: jev-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        classifier_type: llm
        classifier_llm_config:
          model: ollama_chat/qwen2.5:3b
          api_base: http://localhost:11434
          timeout_ms: 2000
        classifier_fallback: heuristic
        tiers:
          SIMPLE: local-llama
          MEDIUM: local-mistral
          COMPLEX: nvidia-deepseek
          REASONING: nvidia-glm
        complexity_router_default_model: local-mistral
</code></pre>
<p>Pull the classifier model first:</p>
<pre><code class="language-shell">ollama pull qwen2.5:3b
</code></pre>
<p><strong>Path 3: Semantic Routing</strong> replaces the classifier with local embeddings.</p>
<pre><code class="language-yaml">model_list:
  - model_name: jev-router
    litellm_params:
      model: auto_router/semantic_router
      semantic_router_config:
        embedding_model: ollama/nomic-embed-text
        routes:
          - name: simple
            utterances:
              - "what is"
              - "define"
              - "hello"
            model: local-llama
          - name: complex
            utterances:
              - "step by step"
              - "analyze this architecture"
              - "refactor the auth module"
            model: nvidia-deepseek
        default_model: local-mistral
</code></pre>
<p>Pull the embedding model first:</p>
<pre><code class="language-shell">ollama pull nomic-embed-text
</code></pre>
<h2>Memori Configuration</h2>
<p>Memori must be configured in <strong>BYODB (Bring Your Own Database)</strong> mode to keep data local. This is done by passing a database connection directly to the <code>Memori</code> constructor. When you provide a valid connection, Memori switches to BYODB mode and sets <code>config.cloud=False</code> internally. Without this connection, it defaults to Memori Cloud.</p>
<h3>Verify SQLite Is Available</h3>
<p>macOS ships with <code>sqlite3</code> pre-installed. Verify it is available:</p>
<pre><code class="language-shell">sqlite3 --version
</code></pre>
<p>Expected output (example):</p>
<pre><code class="language-text">3.43.2 2023-10-10 13:00:00
</code></pre>
<p>For the latest version with extension support, install it via Homebrew:</p>
<pre><code class="language-shell">brew install sqlite
</code></pre>
<p>Homebrew's <code>sqlite</code> is keg-only, so it does not overwrite the system binary. Add it to <code>PATH</code> if you want <code>sqlite3</code> to resolve to the newer version:</p>
<pre><code class="language-shell">echo 'export PATH="/opt/homebrew/opt/sqlite/bin:$PATH"' &gt;&gt; ~/.zshrc
source ~/.zshrc
</code></pre>
<p>The system version at <code>/usr/bin/sqlite3</code> remains available if needed. For the <code>VACUUM;</code> command used below, either version works.</p>
<h3>Create the Local SQLite Database</h3>
<p>The database file must exist before Memori can build its schema inside it. Create the directory and initialize an empty SQLite file:</p>
<pre><code class="language-shell">mkdir -p ~/ai-router/memory
sqlite3 ~/ai-router/memory/memori.db "VACUUM;"
</code></pre>
<p>Expected output: none. Verify the file was created:</p>
<pre><code class="language-shell">ls -lh ~/ai-router/memory/
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">-rw-r--r--  1 YOUR_USERNAME  staff  4.0K Jan  1 12:00 memori.db
</code></pre>
<p>An empty SQLite file is 0 bytes until the first write. The <code>VACUUM;</code> statement forces SQLite to write the file header, so the file appears immediately. If you skip it, the file will be created on the first Memori write.</p>
<p>You can inspect the database at any time:</p>
<pre><code class="language-shell">sqlite3 ~/ai-router/memory/memori.db ".tables"
</code></pre>
<p>Expected output before Memori builds its schema: none. After the schema is built:</p>
<pre><code class="language-text">memori_conversations  memori_memories
</code></pre>
<h3>Use the Database in the Script</h3>
<p>Create <code>~/ai-router/agent.py</code>. The script points Memori at the SQLite file created above, builds the schema, and enables memory:</p>
<pre><code class="language-python">import os
from memori import Memori
from litellm import completion
from sqlalchemy import create_engine
from sqlalchemy.orm import sessionmaker

# 1. Absolute path to the local SQLite database
db_path = os.path.expanduser("~/ai-router/memory/memori.db")

# 2. SQLAlchemy engine and session bound to that file
engine = create_engine(f"sqlite:///{db_path}")
SessionLocal = sessionmaker(bind=engine)

# 3. Initialize Memori in BYODB mode.
# Passing conn=SessionLocal sets config.cloud=False and config.byodb=True
memori = Memori(
    conn=SessionLocal,
    conscious_ingest=True,
)

# 4. Build the schema inside the local database (idempotent)
memori.config.storage.build()

# 5. Enable memory interception for LiteLLM calls
memori.enable()

# 6. Every completion call now reads and writes to the local database
response = completion(
    model="jev-router",
    messages=[{"role": "user", "content": "What did we discuss last time?"}],
    api_base="http://localhost:4000",
    api_key=os.environ["LITELLM_MASTER_KEY"],
)

print(response.choices[0].message.content)
</code></pre>
<p>The key line is <code>conn=SessionLocal</code>. That is what switches Memori out of cloud mode. Without it, the SDK defaults to Memori Cloud and sends data to <code>api.memorilabs.ai</code>.</p>
<h3>Verify the Database Is Being Used</h3>
<p>After running the script once, check that Memori wrote to the local file:</p>
<pre><code class="language-shell">sqlite3 ~/ai-router/memory/memori.db ".tables"
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">memori_conversations  memori_memories
</code></pre>
<p>Count the rows:</p>
<pre><code class="language-shell">sqlite3 ~/ai-router/memory/memori.db \
  "SELECT COUNT(*) FROM memori_conversations;"
</code></pre>
<p>Expected output after one run:</p>
<pre><code class="language-text">1
</code></pre>
<p>If the count is zero, Memori is not writing locally. Confirm the <code>conn</code> parameter is set and the schema build ran without error.</p>
<h3>Verifying Memori Stays Local</h3>
<p>Configuration alone is not a guarantee. The SDK contains code for Memori Cloud endpoints, and a documented "cloud tether" issue may still POST to the cloud even in BYODB mode.</p>
<p><strong>Quick audit with</strong> <code>lsof</code><strong>:</strong></p>
<pre><code class="language-shell">sudo lsof -i -n | grep ESTABLISHED
</code></pre>
<p>Look for entries where the <code>COMMAND</code> is <code>python</code>. If you see connections to IPs that are not your local Ollama (<code>127.0.0.1:11434</code>) or NVIDIA NIM (<code>integrate.api.nvidia.com</code>), that warrants investigation.</p>
<p><strong>Deeper inspection with</strong> <code>tcpdump</code><strong>:</strong></p>
<pre><code class="language-shell">sudo tcpdump -n -i en0 'udp port 53'
</code></pre>
<p>While this runs, execute a test request through the Memori-enabled agent. If you see DNS lookups for domains like <code>api.memorilabs.ai</code>, the tether is active.</p>
<p><strong>Block at the network level:</strong></p>
<pre><code class="language-shell">echo "127.0.0.1 api.memorilabs.ai" | sudo tee -a /etc/hosts
</code></pre>
<p>If the tether is active, the request will fail locally instead of leaving the machine.</p>
<h2>Running as a launchd Service</h2>
<p>The Mac mini runs 24/7. LiteLLM should start on boot and restart if it crashes.</p>
<h3>Create a Wrapper Script</h3>
<p>Instead of putting the entire command and all environment variables directly into the plist, create a small shell script that runs the router. This keeps the plist short and avoids formatting issues.</p>
<p>Create <code>~/ai-router/start_router.sh</code>:</p>
<pre><code class="language-shell">#!/bin/bash
export NVIDIA_NIM_API_KEYS="nvapi-xxxxxxxxxxxxxxxxxxxx"
export LITELLM_MASTER_KEY="sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
export PATH="/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"

cd /Users/YOUR_USERNAME/ai-router
exec /usr/local/bin/litellm --config config.yaml --host 127.0.0.1 --port 4000
</code></pre>
<p>Make it executable:</p>
<pre><code class="language-shell">chmod +x ~/ai-router/start_router.sh
</code></pre>
<h3>Create the launchd Plist</h3>
<p>Create <code>~/Library/LaunchAgents/com.litellm.router.plist</code>:</p>
<pre><code class="language-shell">cat &gt; ~/Library/LaunchAgents/com.litellm.router.plist &lt;&lt; 'EOF'
&lt;?xml version="1.0" encoding="UTF-8"?&gt;
&lt;!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd"&gt;
&lt;plist version="1.0"&gt;
&lt;dict&gt;
    &lt;key&gt;Label&lt;/key&gt;
    &lt;string&gt;com.litellm.router&lt;/string&gt;
    &lt;key&gt;ProgramArguments&lt;/key&gt;
    &lt;array&gt;
        &lt;string&gt;/Users/YOUR_USERNAME/ai-router/start_router.sh&lt;/string&gt;
    &lt;/array&gt;
    &lt;key&gt;RunAtLoad&lt;/key&gt;
    &lt;true/&gt;
    &lt;key&gt;KeepAlive&lt;/key&gt;
    &lt;true/&gt;
    &lt;key&gt;StandardOutPath&lt;/key&gt;
    &lt;string&gt;/Users/YOUR_USERNAME/ai-router/logs/litellm.log&lt;/string&gt;
    &lt;key&gt;StandardErrorPath&lt;/key&gt;
    &lt;string&gt;/Users/YOUR_USERNAME/ai-router/logs/litellm.error.log&lt;/string&gt;
&lt;/dict&gt;
&lt;/plist&gt;
EOF
</code></pre>
<p>Replace <code>YOUR_USERNAME</code> with the macOS short username.</p>
<p>Restrict access to the plist and the wrapper script:</p>
<pre><code class="language-shell">chmod 600 ~/Library/LaunchAgents/com.litellm.router.plist
chmod 700 ~/ai-router/start_router.sh
</code></pre>
<p>Load the service:</p>
<pre><code class="language-shell">launchctl load ~/Library/LaunchAgents/com.litellm.router.plist
</code></pre>
<p>Verify:</p>
<pre><code class="language-shell">launchctl list | grep litellm
</code></pre>
<p>Expected output:</p>
<pre><code class="language-text">12345   0   com.litellm.router
</code></pre>
<p>The first number is the PID. The second is the last exit status. <code>0</code> means success.</p>
<h2>Testing the Full Stack</h2>
<p>Send a request that should route locally:</p>
<pre><code class="language-shell">curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Authorization: Bearer sk-your-virtual-key" \
  -H "Content-Type: application/json" \
  -d '{"model": "jev-router", "messages": [{"role": "user", "content": "Say hello"}]}'
</code></pre>
<p>Expected output (truncated):</p>
<pre><code class="language-json">{
  "model": "local-llama",
  "choices": [{"message": {"content": "Hello! How can I help?"}}]
}
</code></pre>
<p>Confirm the routing decision in the logs:</p>
<pre><code class="language-shell">tail -n 5 ~/ai-router/logs/litellm.log
</code></pre>
<p>Expected output (example):</p>
<pre><code class="language-text">INFO: request model=jev-router, resolved=local-llama, provider=ollama
</code></pre>
<p>Send a request that requires a large context window:</p>
<pre><code class="language-shell">curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Authorization: Bearer sk-your-virtual-key" \
  -H "Content-Type: application/json" \
  -d '{"model": "jev-router", "messages": [{"role": "user", "content": "Analyze this 50-page document..."}]}'
</code></pre>
<p>Expected output (truncated):</p>
<pre><code class="language-json">{
  "model": "nvidia-deepseek",
  "choices": [{"message": {"content": "..."}}]
}
</code></pre>
<p>Confirm the routing decision in the logs:</p>
<pre><code class="language-shell">tail -n 5 ~/ai-router/logs/litellm.log
</code></pre>
<p>Expected output (example):</p>
<pre><code class="language-text">INFO: request model=jev-router, resolved=nvidia-deepseek, provider=nvidia_nim
</code></pre>
<p>The response's <code>model</code> field should report <code>nvidia-deepseek</code> or <code>nvidia-glm</code>, depending on the capability filter.</p>
<h3>Confirming the Routing Rule Is Working</h3>
<p>Send a request that needs vision (an image input). With <code>vision: true</code> only on the NVIDIA models in <code>router.yaml</code>, the router must skip the local models:</p>
<pre><code class="language-shell">curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Authorization: Bearer sk-your-virtual-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-router",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
      ]
    }]
  }'
</code></pre>
<p>The response's <code>model</code> field must be <code>nvidia-deepseek</code> or <code>nvidia-glm</code>. If a local model appears, the <code>vision</code> capability flag in <code>router.yaml</code> is set incorrectly, or the capability check is not detecting the image input.</p>
<h3>Confirming Memori Is Recording</h3>
<p>After a few requests, check the local database:</p>
<pre><code class="language-shell">sqlite3 ~/ai-router/memory/memori.db \
  "SELECT COUNT(*) FROM memori_conversations;"
</code></pre>
<p>The count should increase with each conversation. If it stays at zero, Memori is not intercepting the calls. Confirm <code>memori.enable()</code> ran before the first <code>completion()</code> call.</p>
<h2>Troubleshooting</h2>
<p><code>litellm: command not found</code> — The binary is not in <code>PATH</code>. Use the full path in the plist, or install with <code>pipx</code>.</p>
<p><strong>Service exits immediately after boot</strong> — Check <code>~/ai-router/logs/litellm.error.log</code>. A missing <code>PATH</code> or an unreachable Ollama daemon is the usual cause.</p>
<p><code>ModuleNotFoundError: No module named 'jev_router'</code> — Install both with the same <code>pip3</code>.</p>
<p><strong>401 Unauthorized</strong> — Use a virtual key, not the master key, for client requests.</p>
<p><strong>Requests always route to NVIDIA</strong> — Local models may not be listed in <code>router.yaml</code> or may be priced higher than NVIDIA models.</p>
<p><strong>Memori database not created</strong> — The SQLite path must be absolute. Replace <code>~</code> with the full home directory path.</p>
<p><strong>Memori writes but no rows appear</strong> — Confirm <code>memori.config.storage.build()</code> ran before the first completion call. Check <code>sqlite3 ~/ai-router/memory/memori.db ".tables"</code> for the expected schema.</p>
<p><code>sqlite3: command not found</code> — macOS ships with <code>sqlite3</code> at <code>/usr/bin/sqlite3</code>. If the command is missing, the system installation has been modified. Install via Homebrew with <code>brew install sqlite</code> and add <code>/opt/homebrew/opt/sqlite/bin</code> to <code>PATH</code>.</p>
<p><strong>Suspicious outbound connections from Python</strong> — Use <code>lsof -i -n | grep ESTABLISHED</code> to identify the remote address. If Memori is the source, confirm BYODB mode is active and the <code>/etc/hosts</code> block is in place.</p>
<p><strong>Log file is empty</strong> — Confirm the plist writes to the expected path. The <code>StandardOutPath</code> and <code>StandardErrorPath</code> keys must be absolute paths the launchd user can write to.</p>
<hr />
<p>Happy building.</p>
]]></content:encoded></item><item><title><![CDATA[How to Set Up NVIDIA API Keys and Use Them with Ollama and OpenClaw on a Mac Mini (Part 1)]]></title><description><![CDATA[A Mac mini running Ollama and OpenClaw is a capable local AI stack. The missing piece has always been access to frontier-scale models without needing a GPU cluster. NVIDIA's free API tier on build.nvi]]></description><link>https://nightthoughts.me/how-to-set-up-nvidia-api-keys-and-use-them-with-ollama-and-openclaw-on-a-mac-mini-part-1</link><guid isPermaLink="true">https://nightthoughts.me/how-to-set-up-nvidia-api-keys-and-use-them-with-ollama-and-openclaw-on-a-mac-mini-part-1</guid><category><![CDATA[NVIDIA]]></category><category><![CDATA[llm]]></category><category><![CDATA[ollama]]></category><category><![CDATA[openclaw]]></category><category><![CDATA[Localai]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Mon, 28 Sep 2026 20:36:38 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/8e99da3b-cfd2-4c1a-a7a5-88c6b6b16672.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A Mac mini running Ollama and OpenClaw is a capable local AI stack. The missing piece has always been access to frontier-scale models without needing a GPU cluster. NVIDIA's free API tier on <code>build.nvidia.com</code> fills that gap. It offers open-weight models like DeepSeek, GLM, Kimi, and Nemotron through an OpenAI-compatible endpoint, no credit card required, up to 40 requests per minute.</p>
<p>This guide shows how to obtain an NVIDIA API key, run a local proxy that translates Ollama and OpenAI calls to NVIDIA NIM, and configure OpenClaw to use NVIDIA models directly.</p>
<h2>Before You Start: Ollama and OpenClaw Setup</h2>
<p>This guide assumes Ollama and OpenClaw are already installed, running, and configured on the Mac mini. If they are not, the foundation is covered in the previous guide: <a href="https://nightthoughts.hashnode.dev/i-turned-my-mac-mini-into-a-local-ai-workstation-here-s-exactly-how">I Turned My Mac Mini Into a Local AI Workstation — Here's Exactly How</a>.</p>
<p>That guide covers:</p>
<ul>
<li>Installing Ollama and pulling local models.</li>
<li>Installing OpenClaw and setting up its agentic workflow.</li>
<li>Running both as background services on the Mac mini.</li>
</ul>
<h2>What Is NVIDIA NIM?</h2>
<p>NIM stands for NVIDIA Inference Microservices. It is NVIDIA's packaging of optimized model containers, exposed through a free API at <code>https://integrate.api.nvidia.com/v1</code>. The endpoint follows the OpenAI Chat Completions protocol, so any OpenAI-compatible client can use it after a base-URL change.</p>
<h2>Step 1: Create an NVIDIA API Key</h2>
<ol>
<li>Go to build.nvidia.com and create a free developer account.</li>
<li>Navigate to the API Keys page.</li>
<li>Generate a new key. It will start with <code>nvapi-</code>.</li>
</ol>
<p>On the Mac mini, export the key in the shell profile so it persists across sessions. For Zsh (default on modern macOS):</p>
<pre><code class="language-bash">echo 'export NVIDIA_NIM_API_KEYS="nvapi-xxxxxxxxxxxxxxxxxxxx"' &gt;&gt; ~/.zshrc
source ~/.zshrc
</code></pre>
<p>For Bash:</p>
<pre><code class="language-bash">echo 'export NVIDIA_NIM_API_KEYS="nvapi-xxxxxxxxxxxxxxxxxxxx"' &gt;&gt; ~/.bash_profile
source ~/.bash_profile
</code></pre>
<p>The proxy accepts multiple keys as a comma-separated list. For a single key, the format above is sufficient. To verify the variable is set:</p>
<pre><code class="language-bash">echo $NVIDIA_NIM_API_KEYS
</code></pre>
<blockquote>
<p><strong>Security note:</strong> Never hardcode this key into files committed to Git. Export it in your shell profile instead.</p>
</blockquote>
<h2>Step 2: Run the Ollama Proxy</h2>
<p>Ollama does not natively speak to NVIDIA's cloud endpoint. The <code>jjb8966/ollama-proxy</code> project bridges the gap. It is a Flask-based API gateway that routes requests from Ollama-compatible clients to multiple LLM providers, including NVIDIA NIM, using provider-specific prefixes.</p>
<p>The proxy exposes three API shapes simultaneously: Ollama (<code>/api/chat</code>, <code>/api/tags</code>), OpenAI (<code>/v1/chat/completions</code>, <code>/v1/models</code>), and Anthropic Messages (<code>/v1/messages</code>). Requests are routed to NVIDIA NIM when the model name starts with <code>nvidia-nim:</code>.</p>
<p>This section assumes Ollama is already installed and running as described in the previous Mac mini local AI workstation guide.</p>
<h3>Clone the Repository</h3>
<pre><code class="language-bash">git clone https://github.com/jjb8966/ollama-proxy.git
cd ollama-proxy
</code></pre>
<h3>Install Dependencies</h3>
<p>The project requires Python 3.11 or later.</p>
<pre><code class="language-bash">python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
</code></pre>
<h3>Configure Environment Variables</h3>
<p>The proxy reads two environment variables. Both should be exported in your shell profile, same as the NVIDIA API key.</p>
<ul>
<li><strong><code>NVIDIA_NIM_API_KEYS</code></strong> — already set in Step 1.</li>
<li><strong><code>PROXY_API_TOKEN</code></strong> — a secret token you choose yourself. It is not provided by NVIDIA or the proxy. Clients must present this token in their requests to use the proxy. Think of it as a local password that protects your proxy from unauthorized access on your network. Choose any strong, unique string. You can generate one with <code>openssl rand -hex 32</code>.</li>
</ul>
<p>Add the proxy token to your shell profile. For Zsh:</p>
<pre><code class="language-bash">echo 'export PROXY_API_TOKEN="your-proxy-token-here"' &gt;&gt; ~/.zshrc
source ~/.zshrc
</code></pre>
<p>For Bash:</p>
<pre><code class="language-bash">echo 'export PROXY_API_TOKEN="your-proxy-token-here"' &gt;&gt; ~/.bash_profile
source ~/.bash_profile
</code></pre>
<p>The proxy also supports Google, OpenRouter, Akash, and Cohere, but only the NVIDIA variables are needed for this setup.</p>
<h3>Run the Proxy</h3>
<pre><code class="language-bash">python ollama_proxy.py
</code></pre>
<p>The proxy listens on port <code>5002</code> by default. It can be changed with the <code>PORT</code> environment variable.</p>
<h3>Keep the Proxy Running with launchd</h3>
<p>A Mac mini that runs 24/7 should start the proxy automatically. Create a launchd plist:</p>
<pre><code class="language-bash">cat &gt; ~/Library/LaunchAgents/com.ollama.proxy.plist &lt;&lt; 'EOF'
&lt;?xml version="1.0" encoding="UTF-8"?&gt;
&lt;!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd"&gt;
&lt;plist version="1.0"&gt;
&lt;dict&gt;
    &lt;key&gt;Label&lt;/key&gt;
    &lt;string&gt;com.ollama.proxy&lt;/string&gt;
    &lt;key&gt;WorkingDirectory&lt;/key&gt;
    &lt;string&gt;/Users/YOUR_USERNAME/ollama-proxy&lt;/string&gt;
    &lt;key&gt;ProgramArguments&lt;/key&gt;
    &lt;array&gt;
        &lt;string&gt;/Users/YOUR_USERNAME/ollama-proxy/.venv/bin/python&lt;/string&gt;
        &lt;string&gt;ollama_proxy.py&lt;/string&gt;
    &lt;/array&gt;
    &lt;key&gt;EnvironmentVariables&lt;/key&gt;
    &lt;dict&gt;
        &lt;key&gt;NVIDIA_NIM_API_KEYS&lt;/key&gt;
        &lt;string&gt;nvapi-xxxxxxxxxxxxxxxxxxxx&lt;/string&gt;
        &lt;key&gt;PROXY_API_TOKEN&lt;/key&gt;
        &lt;string&gt;your-proxy-token-here&lt;/string&gt;
        &lt;key&gt;PORT&lt;/key&gt;
        &lt;string&gt;5002&lt;/string&gt;
    &lt;/dict&gt;
    &lt;key&gt;RunAtLoad&lt;/key&gt;
    &lt;true/&gt;
    &lt;key&gt;KeepAlive&lt;/key&gt;
    &lt;true/&gt;
    &lt;key&gt;StandardOutPath&lt;/key&gt;
    &lt;string&gt;/Users/YOUR_USERNAME/Library/Logs/ollama-proxy.log&lt;/string&gt;
    &lt;key&gt;StandardErrorPath&lt;/key&gt;
    &lt;string&gt;/Users/YOUR_USERNAME/Library/Logs/ollama-proxy.error.log&lt;/string&gt;
&lt;/dict&gt;
&lt;/plist&gt;
EOF

launchctl load ~/Library/LaunchAgents/com.ollama.proxy.plist
</code></pre>
<p>Replace <code>YOUR_USERNAME</code> with the macOS short username, found with <code>whoami</code>. Replace the API key and proxy token with the real values.</p>
<blockquote>
<p><strong>Note:</strong> Environment variables must be declared inside the plist. launchd does not read the shell profile, so variables set in <code>~/.zshrc</code> are not available to background services. This is expected and necessary.</p>
</blockquote>
<h3>Use NVIDIA Models Through the Proxy</h3>
<p>Once the proxy is running, clients call it on port <code>5002</code> using the <code>nvidia-nim:</code> prefix. The <code>PROXY_API_TOKEN</code> is already exported in the shell profile, so client code reads it from the environment instead of hardcoding it.</p>
<p>Using the Ollama Python library:</p>
<pre><code class="language-python">import os
import ollama

client = ollama.Client(
    host='http://localhost:5002',
    headers={'Authorization': f"Bearer {os.environ['PROXY_API_TOKEN']}"}
)

response = client.chat(
    model='nvidia-nim:deepseek-ai/deepseek-v4.1-flash',
    messages=[{'role': 'user', 'content': 'Explain MoE in one paragraph.'}]
)

print(response['message']['content'])
</code></pre>
<p>Expected output (example):</p>
<p>Mixture of Experts (MoE) is a technique where multiple specialized sub-models, called experts, are combined. A gating network routes each input token to only a small subset of experts, so the model can have a very large number of parameters while keeping computation low. This makes MoE models efficient and scalable for large language tasks.</p>
<p>Or through the OpenAI-compatible route:</p>
<pre><code class="language-python">import os
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:5002/v1",
    api_key=os.environ["PROXY_API_TOKEN"]
)

completion = client.chat.completions.create(
    model="nvidia-nim:z-ai/glm-5-3",
    messages=[{"role": "user", "content": "Write a haiku about GPUs."}]
)

print(completion.choices[0].message.content)
</code></pre>
<p>Expected output (example):</p>
<p>Silicon chips glow,
Parallel cores hum softly,
Numbers dance in light.</p>
<p>List all available models with:</p>
<pre><code class="language-bash">curl -H "Authorization: Bearer $PROXY_API_TOKEN" \
  http://localhost:5002/api/tags
</code></pre>
<p>Useful model IDs include:</p>
<ul>
<li><code>nvidia-nim:deepseek-ai/deepseek-v4.1-flash</code></li>
<li><code>nvidia-nim:z-ai/glm-5-3</code></li>
<li><code>nvidia-nim:z-ai/glm-5-3-flash</code></li>
<li><code>nvidia-nim:moonshotai/kimi-k3</code></li>
<li><code>nvidia-nim:nvidia/nemotron-3-ultra-550b-a55b</code></li>
</ul>
<blockquote>
<p>Model IDs follow NVIDIA's catalog naming. Confirm availability with <code>GET /v1/models</code> on the endpoint if a call returns <code>404</code>. The proxy's model list is defined in its <code>models.json</code> file.</p>
</blockquote>
<h2>Why Use a Proxy Alongside OpenClaw?</h2>
<p>OpenClaw can talk to NVIDIA directly. So why add a proxy at all? Because the Mac mini is more than just OpenClaw.</p>
<p>Ollama is the home base for local models on the machine. Many tools already know how to talk to Ollama. They use its API. They do not know about NVIDIA. Changing every tool to support NVIDIA separately would be a lot of work.</p>
<p>The proxy solves that. It sits in front of NVIDIA and looks like Ollama to the rest of the system. Any tool that already uses Ollama can now use NVIDIA models by changing only the model name. No new SDK. No new provider code.</p>
<p>This keeps one simple setup:</p>
<ul>
<li>Local models for quick tasks, private work, and offline use.</li>
<li>NVIDIA models for heavy reasoning, long context, and multimodal jobs.</li>
<li>Same API for both.</li>
</ul>
<p>The proxy also keeps the NVIDIA API key in one place. Tools do not each need their own key. That is easier to manage and safer.</p>
<p>OpenClaw still uses its native NVIDIA provider for the best agentic experience. The proxy is for everything else. If OpenClaw is the only tool running, the proxy is not needed. If other Ollama-based tools are running, the proxy makes them cloud-capable with almost no effort.</p>
<p>The previous Mac mini local AI workstation guide already sets Ollama as the central hub. The proxy extends that hub to cloud models without disrupting the rest of the setup.</p>
<h2>Step 3: Integrate NVIDIA with OpenClaw</h2>
<p>OpenClaw natively supports NVIDIA as a model provider. It auto-enables when the <code>NVIDIA_API_KEY</code> environment variable is set and defaults to the Nemotron 3 Ultra model.</p>
<p>This section assumes OpenClaw is already installed and configured as described in the previous Mac mini local AI workstation guide. The NVIDIA API key was already exported in Step 1, though OpenClaw expects it under a different variable name. Set it:</p>
<pre><code class="language-bash">echo 'export NVIDIA_API_KEY="nvapi-xxxxxxxxxxxxxxxxxxxx"' &gt;&gt; ~/.zshrc
source ~/.zshrc
</code></pre>
<p>Then run the onboarding command:</p>
<pre><code class="language-bash">openclaw onboard --auth-choice nvidia-api-key
</code></pre>
<p>Then set the default model:</p>
<pre><code class="language-bash">openclaw models set nvidia/nvidia/nemotron-3-ultra-550b-a55b
</code></pre>
<p>If OpenClaw runs as a background service on the Mac mini, ensure the environment variable is available to the service. If set in <code>~/.zshrc</code>, that only applies to interactive shells. For launchd services, add it to the plist's <code>EnvironmentVariables</code> section, similar to the proxy setup above.</p>
<h3>Manual Config Snippet</h3>
<p>To edit the OpenClaw config manually, use this YAML structure:</p>
<pre><code class="language-yaml">env:
  NVIDIA_API_KEY: "nvapi-xxxxxxxxxxxxxxxxxxxx"

models:
  providers:
    nvidia:
      baseUrl: "https://integrate.api.nvidia.com/v1"
      api: "openai-completions"

agents:
  defaults:
    model:
      primary: "nvidia/nvidia/nemotron-3-ultra-550b-a55b"
</code></pre>
<p>This config auto-loads NVIDIA's featured model catalog from <code>assets.ngc.nvidia.com</code> and caches it for 24 hours, so new models appear without an OpenClaw update.</p>
<h3>Use NVIDIA Models in OpenClaw</h3>
<p>Switch models interactively within OpenClaw:</p>
<pre><code>/model nvidia/nvidia/nemotron-3-ultra-550b-a55b
</code></pre>
<p>Or set it as the default for agentic workflows:</p>
<pre><code class="language-bash">openclaw agents set-default-model nvidia/nvidia/nemotron-3-ultra-550b-a55b
</code></pre>
<p>Nemotron 3 Ultra is a 550B total parameter model with 55B active and a 1M-token context window. It is built for long-context agentic work, which suits OpenClaw's multi-step reasoning and tool-calling tasks. Lighter alternatives in the built-in fallback catalog include:</p>
<ul>
<li><code>nvidia/nvidia/nemotron-3-super-120b</code></li>
<li><code>nvidia/nvidia/llama-3.1-nemotron-70b-instruct</code></li>
<li><code>meta/llama-3.3-70b-instruct</code></li>
<li><code>nvidia/mistral-nemo-minitron-8b-8k-instruct</code></li>
</ul>
<h2>The Hybrid Stack</h2>
<p>With these changes, the Mac mini runs a hybrid AI stack:</p>
<ol>
<li>NVIDIA API key is exported in the shell profile and available to background services.</li>
<li>Ollama proxy runs on port <code>5002</code> via launchd, providing access to NVIDIA NIM models for Ollama-based tools.</li>
<li>OpenClaw uses NVIDIA NIM directly for agentic coding sessions, with Nemotron 3 Ultra as the default model.</li>
<li>Local Ollama models remain available on <code>localhost:11434</code> for quick tasks, privacy-sensitive work, and offline use.</li>
</ol>
<p>Local models handle speed and privacy. Cloud models handle scale and capability. These two paths are separate for now.</p>
<h2>Rate Limits and Scaling</h2>
<p>The free tier allows 40 requests per minute for most models. For prototyping and individual use, that is sufficient. If the limit is reached, options include:</p>
<ol>
<li>Wait and retry — the limit resets quickly.</li>
<li>Deploy a NIM container locally on a GPU-equipped machine.</li>
<li>Use a partner endpoint like AWS, Azure, or GCP, which host NIM containers.</li>
</ol>
<h2>Troubleshooting</h2>
<p><strong>"Connection refused" on the proxy:</strong> Ensure the proxy is running (<code>launchctl list | grep ollama</code>) and that <code>NVIDIA_NIM_API_KEYS</code> is set in the same shell session.</p>
<p><strong>401 Unauthorized from the proxy:</strong> The <code>Authorization: Bearer</code> header is missing or the <code>PROXY_API_TOKEN</code> does not match. Include the header in every client request.</p>
<p><strong>401 Unauthorized from NVIDIA:</strong> The NVIDIA API key may have expired or been revoked. Generate a new one at <code>build.nvidia.com</code>.</p>
<p><strong>404 Model not found:</strong> Check the exact model ID against the proxy's <code>models.json</code> file or NVIDIA's catalog with <code>GET /v1/models</code>. The <code>nvidia-nim:</code> prefix is required for the proxy, and the <code>nvidia/</code> prefix is required for OpenClaw model references.</p>
<p><strong>Proxy exits immediately after boot:</strong> Check <code>~/Library/Logs/ollama-proxy.error.log</code>. A wrong interpreter path or a missing <code>WorkingDirectory</code> is the usual cause.</p>
<hr />
<p>Happy building.</p>
]]></content:encoded></item><item><title><![CDATA[NVIDIA Is Giving Away Access to Powerful AI Models for Free]]></title><description><![CDATA[Imagine you want to use a smart AI model to help you write code, summarize a long document, or answer questions. Normally, you either need a very expensive computer or you pay a company for every time]]></description><link>https://nightthoughts.me/nvidia-is-giving-away-access-to-powerful-ai-models-for-free</link><guid isPermaLink="true">https://nightthoughts.me/nvidia-is-giving-away-access-to-powerful-ai-models-for-free</guid><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[NVIDIA]]></category><category><![CDATA[open source]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 27 Sep 2026 15:35:15 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/87b6eb16-23d5-4523-a6b2-8bebde5bb823.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine you want to use a smart AI model to help you write code, summarize a long document, or answer questions. Normally, you either need a very expensive computer or you pay a company for every time you use their AI.</p>
<p>NVIDIA is now offering a third option: you can use some of the world's most powerful AI models for free through their online API. No expensive hardware, no monthly subscription. All you need is an internet connection and a free account.</p>
<h2>The Free Models You Can Use</h2>
<p>NVIDIA has opened free access to a catalog of open-weight AI models through its developer platform at <strong>build.nvidia.com</strong>. This means you can use these models without paying for tokens or requests. Some of the models available include:</p>
<ul>
<li><p><strong>DeepSeek V4.1 Flash</strong> – good for fast reasoning and coding tasks.</p>
</li>
<li><p><strong>GLM 5.3</strong> – a very large model designed for long coding sessions and tool use.</p>
</li>
<li><p><strong>GLM 5.3 Flash</strong> – a faster, lighter version of GLM.</p>
</li>
<li><p><strong>Kimi K3</strong> – a massive model built for long conversations and coding.</p>
</li>
</ul>
<p>There are also other models like Nemotron and MiniMax, and the catalog keeps growing.</p>
<p>These are not small, stripped-down demo versions. Kimi K3, for example, is a huge model with billions of parameters. Running it on your own computer would require a server room full of expensive hardware. NVIDIA hosts everything on their own powerful machines.</p>
<h2>How to Get Started</h2>
<p>The process is simple and free. Here is what you do:</p>
<ol>
<li><p><strong>Create a free account</strong> at <a href="https://build.nvidia.com">build.nvidia.com</a>.</p>
</li>
<li><p><strong>Find a model</strong> with the "Free Endpoint" label. These are the ones you can use without paying.</p>
</li>
<li><p><strong>Generate an API key</strong> by clicking your profile icon, going to API Keys, and creating a new key.</p>
</li>
<li><p><strong>Start sending requests</strong> to the API.</p>
</li>
</ol>
<p>The API is <strong>OpenAI-compatible</strong>, which means if you already know how to use OpenAI's API, you can use NVIDIA's the same way. You just change the base URL to:</p>
<pre><code class="language-text">https://integrate.api.nvidia.com/v1
</code></pre>
<p>And you use your NVIDIA API key instead of an OpenAI key.</p>
<h2>A Simple Example</h2>
<p>Let's say you want to ask DeepSeek V4.1 Flash a question. You would send a request like this:</p>
<pre><code class="language-bash">curl --request POST \
  --url https://integrate.api.nvidia.com/v1/chat/completions \
  --header 'Authorization: Bearer YOUR_NVIDIA_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "deepseek-ai/deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Write a short poem about the ocean."}],
    "max_tokens": 100
  }'
</code></pre>
<p>The model will send back a response. You can do this from any programming language, or even from a tool like Postman.</p>
<h2>What You Can Do With It</h2>
<p>Here are some practical things you can build:</p>
<ul>
<li><p><strong>A coding helper</strong> that suggests fixes for your code.</p>
</li>
<li><p><strong>A document summarizer</strong> that reads long articles and gives you the key points.</p>
</li>
<li><p><strong>A chatbot</strong> for your website that answers common questions.</p>
</li>
<li><p><strong>A study assistant</strong> that explains difficult topics in simple terms.</p>
</li>
</ul>
<p>You can also connect these models to coding tools like Cursor or OpenCode, which have built-in NVIDIA integration.</p>
<h2>A Few Things to Know</h2>
<p>The free tier runs on a <strong>rate limit of about 40 requests per minute</strong> for most models. This means you can make up to 40 requests every minute. For personal projects, learning, and small experiments, this is usually more than enough.</p>
<p>If you need more, you can always look at the "Deploy" section on NVIDIA's website to run the model on your own infrastructure or through a cloud partner.</p>
<p>There is also a <strong>free credit allowance</strong> for new accounts, which gives you some extra room before you ever need to think about limits.</p>
<h2>Final Thoughts</h2>
<p>NVIDIA's free API catalog is a practical way to experiment with powerful AI models without spending money or buying hardware. It is not a fully managed service, and the free tier has limits, but for learning and building small projects, it is a great starting point.</p>
<p>If you have been curious about trying AI models in your own projects, this is a good moment to start.</p>
<p><strong>Try it here:</strong> <a href="https://build.nvidia.com">build.nvidia.com</a></p>
<p><strong>API Base URL:</strong> <code>https://integrate.api.nvidia.com/v1</code></p>
]]></content:encoded></item><item><title><![CDATA[Buzz: The Open-Source Workspace Where AI Joins Your Team]]></title><description><![CDATA[Have you ever worked on a group project? You have a group chat, you share files, and everyone has a role. Now imagine one of your teammates is an AI. It can read the chat, help with code, and do tasks]]></description><link>https://nightthoughts.me/buzz-the-open-source-workspace-where-ai-joins-your-team</link><guid isPermaLink="true">https://nightthoughts.me/buzz-the-open-source-workspace-where-ai-joins-your-team</guid><category><![CDATA[AI]]></category><category><![CDATA[Collaboration]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[ai agents]]></category><category><![CDATA[Developer Tools]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 27 Sep 2026 15:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/03cbc45b-1838-43d1-a614-9a76ce8e2c72.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Have you ever worked on a group project? You have a group chat, you share files, and everyone has a role. Now imagine one of your teammates is an AI. It can read the chat, help with code, and do tasks. That is the idea behind <strong>Buzz</strong>.</p>
<p>Buzz is an open-source project from Jack Dorsey’s company, <strong>Block</strong>. It is not a magic robot that runs a whole company. It is a tool that lets people and AI agents work together in one place.</p>
<p>An <strong>AI agent</strong> is a computer program that can do tasks on its own—like reading code, writing a message, or opening a file.</p>
<h2>What is Buzz?</h2>
<p>Think of it as <strong>Slack + GitHub + AI helpers</strong> in one app. You can:</p>
<ul>
<li><p>Chat with your team.</p>
</li>
<li><p>Share code and files.</p>
</li>
<li><p>Ask AI agents to do work.</p>
</li>
<li><p>See everything in one history.</p>
</li>
</ul>
<p>It uses <strong>Nostr</strong>, which is a way to give every person and every AI its own digital ID. That ID is like a username that cannot be faked.</p>
<h2>A Simple Example</h2>
<p>You make a channel called <code>website-fix</code>. You invite:</p>
<ul>
<li><p>Your friend Maria (a human).</p>
</li>
<li><p>An AI agent named <code>Fixer</code>.</p>
</li>
</ul>
<p>You write: “The login button is broken.”</p>
<p><code>Fixer</code> can look at the code, suggest a fix, and open a pull request. Maria can review it. All of this happens in the same chat. You do not need to jump between five different apps.</p>
<h2>Key Benefits</h2>
<ol>
<li><p><strong>AI gets its own account.</strong><br />It is not just a bot you command. It has a name and an ID. You can see what it did and when.</p>
</li>
<li><p><strong>One place for everything.</strong><br />Chat, code, tasks, and AI work live together. This saves time.</p>
</li>
<li><p><strong>Works with different AI models.</strong><br />You can use Claude, ChatGPT, or other AI tools. You are not locked into one company.</p>
</li>
<li><p><strong>You can run it yourself.</strong><br />Because it is open source, you can host it on your own computer or server. You control your data.</p>
</li>
<li><p><strong>Clear history.</strong><br />Every action is saved. It is like a group chat log, but it also includes code changes.</p>
</li>
</ol>
<h2>What Buzz Is Not</h2>
<ul>
<li><p>It is <strong>not</strong> a fully automatic company. Humans still need to lead and make decisions.</p>
</li>
<li><p>It is <strong>not</strong> a replacement for all your current tools yet. It is early.</p>
</li>
<li><p>It is <strong>not</strong> perfect. The project is still growing.</p>
</li>
</ul>
<p>Some viral posts say it can run a company “100% by AI.” That is not what it is. It is a workspace for humans and AI to work together.</p>
<h2>How to Try It</h2>
<p>Go to the GitHub page and read the README:</p>
<p><strong><a href="https://github.com/block/buzz">https://github.com/block/buzz</a></strong></p>
<p>You can download the app for Mac, Windows, or Linux. Start small: create one channel and invite one AI agent. See how it feels.</p>
<h2>Final Thoughts</h2>
<p>Buzz is interesting because it treats AI like a teammate, not just a tool. It is a way to explore how humans and AI can work together in one shared space. Ignore the hype about “100% AI companies.” The real value is collaboration.</p>
]]></content:encoded></item><item><title><![CDATA[Security and Governance: The Guardrails That Make AI Safe to Use]]></title><description><![CDATA[Artificial intelligence offers seemingly endless possibilities. It can write, design, diagnose, forecast, negotiate, and increasingly act on our behalf. The potential feels unlimited: faster decisions]]></description><link>https://nightthoughts.me/security-and-governance-the-guardrails-that-make-ai-safe-to-use</link><guid isPermaLink="true">https://nightthoughts.me/security-and-governance-the-guardrails-that-make-ai-safe-to-use</guid><category><![CDATA[AI]]></category><category><![CDATA[Security]]></category><category><![CDATA[Governance]]></category><category><![CDATA[cybersecurity]]></category><category><![CDATA[Machine Learning]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Wed, 23 Sep 2026 16:17:12 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/4ac6066d-a350-4e82-ad16-d9a1ceede9a2.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Artificial intelligence offers seemingly endless possibilities. It can write, design, diagnose, forecast, negotiate, and increasingly act on our behalf. The potential feels unlimited: faster decisions, lower costs, new services, and solutions to problems once thought unsolvable.</p>
<p>But as AI moves from answering questions to taking action, that potential introduces a new kind of risk. A bad output is inconvenient. A bad action can cost money, expose data, disrupt operations, or cause harm. The question is no longer only <em>"What can AI do?"</em> but <em>"What should AI be allowed to do, and how do we keep it safe?"</em></p>
<p>The answer lies in two complementary ideas: <strong>security</strong> and <strong>governance</strong>. Together, they form the guardrails that make it possible to move quickly without crashing.</p>
<hr />
<h2>What Are AI Guardrails?</h2>
<p>A guardrail does not stop a car from moving. It keeps the car on the road when something goes wrong. The same applies to AI.</p>
<p>AI guardrails are the combination of technical controls and organizational rules that keep AI systems safe, accountable, and aligned with human intent. They rest on two pillars.</p>
<p><strong>Security: the lock.</strong> Security asks: <em>Can someone misuse this?</em> It protects the AI system from being manipulated, attacked, or broken—and protects people and organizations from the consequences when AI fails.</p>
<p><strong>Governance: the rulebook.</strong> Governance asks: <em>Should this be allowed, by whom, and how do we know?</em> It covers policies, approvals, human oversight, audit logs, and accountability.</p>
<p>Neither works alone. Security without governance is a locked door with no rules about who gets a key. Governance without security is a rulebook sitting next to an open vault.</p>
<table>
<thead>
<tr>
<th></th>
<th>Security</th>
<th>Governance</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Focus</strong></td>
<td>Protection</td>
<td>Direction</td>
</tr>
<tr>
<td><strong>Question</strong></td>
<td>Can someone misuse this?</td>
<td>Should this be allowed, by whom, and how do we know?</td>
</tr>
<tr>
<td><strong>Analogy</strong></td>
<td>The lock and alarm</td>
<td>The rulebook and referee</td>
</tr>
<tr>
<td><strong>Examples</strong></td>
<td>Access controls, encryption, monitoring</td>
<td>Policies, approvals, audit logs, oversight</td>
</tr>
</tbody></table>
<p>Guardrails are not barriers to innovation. They are what make innovation sustainable, because they allow AI to be trusted with greater responsibility.</p>
<hr />
<h2>Security: Protecting AI and Protecting from AI</h2>
<p>Security focuses on one question: <em>Can someone misuse this?</em> As AI gains access to databases, money, and machines, the attack surface expands. Common risks include:</p>
<ul>
<li><p><strong>Prompt injection</strong> — tricking the AI into ignoring its instructions.</p>
</li>
<li><p><strong>Data leakage</strong> — exposing private or sensitive information.</p>
</li>
<li><p><strong>Unauthorized access</strong> — gaining entry to the AI system or its data.</p>
</li>
<li><p><strong>Model theft</strong> — stealing the AI model itself.</p>
</li>
<li><p><strong>Supply chain vulnerabilities</strong> — compromised third-party tools, libraries, or data.</p>
</li>
</ul>
<p>Security controls reduce these risks. Access controls ensure only authorized users and systems can interact with the AI. Encryption protects data at rest and in transit. Least privilege limits the AI to the minimum permissions it needs. Monitoring and logging record important actions for review. Regular security testing finds vulnerabilities before attackers do.</p>
<p>Security is the lock on the door and the alarm on the wall. It does not decide what the AI should do. It ensures only the right actors get in, and that anything unusual is detected and recorded.</p>
<hr />
<h2>Governance: Rules, Roles, and Accountability</h2>
<p>Governance decides what the system is allowed to do. It answers: <em>Should this be allowed, by whom, and how do we know?</em></p>
<p>An AI system can be perfectly secure and still cause harm. It can make unfair decisions, violate regulations, or act beyond its scope. Governance provides direction, boundaries, and oversight.</p>
<p>Core elements include:</p>
<ul>
<li><p><strong>Policies and acceptable use</strong> — written rules about what AI can and cannot be used for.</p>
</li>
<li><p><strong>Approval workflows</strong> — defined steps for who approves an AI system and which actions require sign-off.</p>
</li>
<li><p><strong>Human-in-the-loop</strong> — a human reviews and confirms the AI's recommendation before high-consequence decisions take effect.</p>
</li>
<li><p><strong>Audit logs and traceability</strong> — every important action is recorded, enabling investigation and accountability.</p>
</li>
<li><p><strong>Risk classification</strong> — systems are sorted by risk level, with stricter controls applied where stakes are higher.</p>
</li>
<li><p><strong>Compliance and ethics</strong> — ensuring AI meets legal, regulatory, and ethical standards, including privacy, fairness, and transparency.</p>
</li>
</ul>
<p>Governance is the rulebook and the referee. It does not play the game. It makes sure the game is played fairly and according to agreed rules.</p>
<hr />
<h2>How Security and Governance Work Together</h2>
<p>Each pillar covers what the other cannot.</p>
<p>Security without governance is a hardened system with no policy defining what it should or should not do. It remains protected from outsiders, but it can still make harmful or non-compliant decisions.</p>
<p>Governance without security is a detailed rulebook with no enforcement. The rules exist, but the system can be tricked or hacked into ignoring them.</p>
<p>Together, they create <strong>safe autonomy</strong>: an AI system that is protected from misuse and operates within clearly defined boundaries. This is the foundation of trust, and trust is what allows AI to move from controlled experiments to real-world impact.</p>
<hr />
<h2>Real-World Applications</h2>
<p><strong>Customer service AI</strong> answers questions, processes refunds, and escalates complex issues. Security protects against prompt injection and unauthorized access to customer records. Governance sets refund limits and requires human approval above a threshold. Every action is logged.</p>
<p><strong>Healthcare AI</strong> assists with diagnosis and treatment recommendations. Security protects patient data and restricts access to authorized staff. Governance requires a licensed clinician to approve any AI-generated diagnosis or treatment plan, with audit trails recording who saw what and when.</p>
<p><strong>Self-driving systems</strong> control steering, acceleration, and braking. Security protects vehicle software from being hacked and sensor data from being spoofed. Governance sets safety standards, operational limits, and incident reporting requirements, with regulations defining responsibility when something goes wrong.</p>
<p><strong>Enterprise AI agents</strong> automate workflows such as scheduling, expense approval, and supply chain management. Security applies least privilege and monitors for unusual activity. Governance defines spending limits, approval chains, and which actions require human confirmation.</p>
<p><strong>Finance and banking</strong> uses AI for fraud detection, credit scoring, loan approvals, wealth management, and customer service. It can also act as an agent that moves money or rebalances portfolios. Security protects against data breaches, unauthorized transactions, and model theft. Governance is especially strict here: credit-scoring and loan-approval AI is classified as high-risk under regulations like the EU AI Act, making human oversight, data governance, and explainability mandatory. For AI agents that move money, governance sets predefined mandates, and every proposed action is validated and recorded before it executes. Audit logs are required for compliance. The AI can act, but only within approved, monitored, and recorded rules.</p>
<p>In every case, security protects the system and the data. Governance defines what the AI may do, who approves it, and how we know what happened.</p>
<hr />
<h2>Frameworks for AI Security and Governance</h2>
<p>Frameworks are structured guides—recipes, not rigid rules—that organizations can adapt to their own needs.</p>
<p><strong>NIST AI Risk Management Framework (AI RMF)</strong> is a voluntary framework built around four functions: <strong>Govern</strong> (establish roles and policies), <strong>Map</strong> (understand where AI is used and what could go wrong), <strong>Measure</strong> (assess risks consistently), and <strong>Manage</strong> (reduce risks and monitor results). It also defines what trustworthy AI looks like: valid, safe, secure, accountable, transparent, explainable, privacy-enhanced, and fair. It is free, sector-agnostic, and widely used as a starting point.</p>
<p><strong>ISO/IEC 42001</strong> is the first international certifiable standard for an AI Management System. Where NIST offers guidance, ISO/IEC 42001 allows formal certification, demonstrating to customers, regulators, and partners that AI is managed responsibly. It applies to any organization that builds, buys, or uses AI.</p>
<p><strong>OWASP GenAI Security Project</strong> maintains the <strong>Top 10 for LLM Applications</strong>, updated annually based on real-world incident data. The 2026 edition ranks prompt injection, sensitive information disclosure, and excessive agency as the top three risks. A companion <strong>Top 10 for Agentic Applications</strong> focuses on what autonomous systems are allowed to do once they move from reasoning to action.</p>
<p><strong>EU AI Act</strong> is a binding regulation that classifies AI systems by risk level. Creditworthiness and credit-scoring systems are explicitly high-risk, triggering requirements for human oversight, data governance, transparency, and conformity assessment.</p>
<p><strong>MAS SAFR (Safeguards for Agentic Finance at Runtime)</strong> was developed by the Monetary Authority of Singapore for AI agents in financial services. It introduces governance checkpoints that verify and record an agent's proposed actions before they execute, keeping behavior within predefined mandates and risk boundaries. It has been applied to agent-assisted payments, wealth management, and client engagement.</p>
<p>These frameworks differ in scope and formality, but they share a core idea: AI should operate within clearly defined boundaries, with oversight, transparency, and accountability. Start with one and build from there.</p>
<hr />
<h2>Practical Checklist</h2>
<p><strong>Inventory your AI tools and data access.</strong> Know what systems are in use, what data they can reach, and what actions they can take.</p>
<p><strong>Apply least privilege.</strong> Give each system only the minimum permissions it needs.</p>
<p><strong>Log all important actions.</strong> Record what the AI does, when, and who or what initiated it.</p>
<p><strong>Require human approval for high-risk tasks.</strong> Define which actions are too consequential to automate fully.</p>
<p><strong>Test for prompt injection and data leakage.</strong> Probe your systems regularly and fix what you find.</p>
<p><strong>Write a simple AI use policy.</strong> Document what is allowed, what is prohibited, who approves new use cases, and how incidents are reported.</p>
<p><strong>Review and update regularly.</strong> Capabilities, risks, and regulations evolve quickly. What was safe six months ago may not be safe today.</p>
<p>Security and governance are ongoing practices, not one-time projects.</p>
<hr />
<h2>Conclusion</h2>
<p>AI offers seemingly endless possibilities, and those possibilities are growing. But as AI moves from answering questions to taking action, the risks grow with it.</p>
<p>Security and governance are the guardrails that make it possible to explore those possibilities safely. Security protects AI systems from being hacked, tricked, or misused. Governance defines what AI is allowed to do, who approves it, and how we know what happened. Together, they create safe autonomy—the ability to trust AI with greater responsibility because the boundaries are clear and the protections are in place.</p>
<p>Frameworks like the NIST AI RMF, ISO/IEC 42001, and the OWASP Top 10 for LLM Applications provide practical blueprints. Real-world examples from customer service, healthcare, self-driving systems, enterprise agents, and finance show that guardrails are already being applied in production systems today.</p>
<p>Trust leads to adoption. Adoption leads to value. Guardrails are what make that journey possible.</p>
]]></content:encoded></item><item><title><![CDATA[Fast Brain, Slow Brain: A Beginner's Guide to System 1 and System 2 AI]]></title><description><![CDATA[You ask an AI "is this email spam?" and it writes you a 200-word essay explaining its reasoning, citing signals, and hedging its conclusion.
Your brain doesn't work that way. You'd just know.
In Septe]]></description><link>https://nightthoughts.me/fast-brain-slow-brain-a-beginner-s-guide-to-system-1-and-system-2-ai</link><guid isPermaLink="true">https://nightthoughts.me/fast-brain-slow-brain-a-beginner-s-guide-to-system-1-and-system-2-ai</guid><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[jev]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Wed, 23 Sep 2026 05:41:50 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/031588b9-1bd9-4e63-a9c7-7fde47ce28f2.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You ask an AI "is this email spam?" and it writes you a 200-word essay explaining its reasoning, citing signals, and hedging its conclusion.</p>
<p>Your brain doesn't work that way. You'd just <em>know</em>.</p>
<p>In September 2026, a model called Jev hit 140,000 waitlist signups in 36 hours by doing exactly that: no essays, just decisions. This article explains what it is, why it matters, and when you'd actually use it.</p>
<hr />
<h2>The Two Systems of the Mind</h2>
<p>In his book <em>Thinking, Fast and Slow</em>, psychologist Daniel Kahneman described two modes of thought.</p>
<p><strong>System 1</strong> is fast, automatic, and silent. You're walking on a trail and see a curved shape in the grass. Before you can form the word "snake," your body has already flinched. No reasoning. No explanation. Just a decision.</p>
<p><strong>System 2</strong> is slow, deliberate, and verbal. You're doing your taxes. You read the instructions, consider each deduction, check your arithmetic. It's effortful and it feels like thinking.</p>
<p>Here's the mapping that matters for AI:</p>
<ul>
<li><p><strong>Today's famous models</strong> — ChatGPT, Claude, Gemini — are artificial System 2. They think by writing, one token at a time. That's why they're brilliant at reasoning and terrible at reflexes.</p>
</li>
<li><p><strong>A new category of models</strong> aims to be artificial System 1: instant, silent, typed decisions.</p>
</li>
</ul>
<hr />
<h2>Why This Matters Now</h2>
<p>Most AI calls in production aren't creative questions. They're judgments:</p>
<ul>
<li><p>"Which team should handle this support ticket?"</p>
</li>
<li><p>"Is this transaction risky?"</p>
</li>
<li><p>"Should this agent be allowed to run this tool?"</p>
</li>
</ul>
<p>Right now, we pay System 2 prices — 10 to 40 seconds and cents per call — for System 1 jobs that should take milliseconds and fractions of a cent.</p>
<p>To be fair, classifiers and rerankers have existed for years. What's new is packaging a frontier-quality decision model as a hosted API that anyone can call with a credit card.</p>
<p>And the adoption has been fast. When Jev launched on September 15, 2026, roughly 13% of Vercel's paid AI Gateway teams were using it within 24 hours — more than double any previous model launch.</p>
<hr />
<h2>What Is a System One Model?</h2>
<p>Simple definition: unstructured input goes in, typed probabilistic decisions come out. No text generated.</p>
<p>Think of a restaurant host. A party walks in. The host doesn't write an essay about seating theory. They glance at the room and make a decision — instantly, silently, and in one of three shapes:</p>
<table>
<thead>
<tr>
<th>Decision shape</th>
<th>What the host does</th>
<th>What it returns</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Choice</strong></td>
<td>Picks one option from a fixed list</td>
<td>"Booth section"</td>
</tr>
<tr>
<td><strong>Score</strong></td>
<td>Rates something on a scale</td>
<td>"Wait time: 0.7 out of 1"</td>
</tr>
<tr>
<td><strong>Noul</strong></td>
<td>Answers a yes/no question with confidence</td>
<td>"Reservation on file: yes, 0.95"</td>
</tr>
</tbody></table>
<p>That's the whole idea. The host never explains. They never write a paragraph. They just return a typed answer your system can act on directly.</p>
<p>Yes, LLMs can also output JSON. System One models are different in <em>what they're trained to do</em> — calibrated decisions, not just formatted text. Vercel's own comparison notes that Jev "evaluates each question independently against the same state, returning typed answers with probabilities over the defined outcomes," while LLM probabilities are "generated estimates."</p>
<p>The first public example is Jev by TypeSafe AI. You can call it on OpenRouter with the model ID <code>typesafe/jev-1.13</code>, or through Vercel AI Gateway as <code>typesafe-ai/jev</code>.</p>
<hr />
<h2>So How Much Faster and Cheaper, Really?</h2>
<p>Independent testers ran Jev against a comparable LLM on the same decision tasks:</p>
<table>
<thead>
<tr>
<th>Metric</th>
<th>Jev</th>
<th>Comparable LLM</th>
</tr>
</thead>
<tbody><tr>
<td>Median latency</td>
<td>105 ms</td>
<td>710 ms</td>
</tr>
<tr>
<td>Cost per 1,000 decisions</td>
<td>$0.04</td>
<td>$0.16</td>
</tr>
<tr>
<td>Accuracy (vendor dashboard)</td>
<td>67.8%</td>
<td>74.1%</td>
</tr>
</tbody></table>
<p>Read that table carefully. Jev is faster and cheaper by an order of magnitude — and slightly <em>less</em> accurate. TypeSafe's own claims of 20–200x speed and 40–400x cost reduction are self-tested against agreement with other frontier models rather than ground truth, and the one independent check (Every) found roughly 25x faster and 580x cheaper on extraction tasks specifically — "good but not perfect."</p>
<p>So here's the honest trade-off. Jev isn't more accurate than the best large models. It's slightly less accurate, and dramatically cheaper and faster. For most production decisions, that's the right trade. For a legal contract or a medical diagnosis, it isn't.</p>
<p>The framing to remember: you don't hire a decision model because it's smarter. You hire it because most decisions don't need smart. They need fast, cheap, and consistent.</p>
<hr />
<h2>Which Jobs for Which Brain?</h2>
<table>
<thead>
<tr>
<th>Task</th>
<th>System One (decide fast)</th>
<th>System Two (think in words)</th>
</tr>
</thead>
<tbody><tr>
<td>"Is this ticket urgent?"</td>
<td>✅ ideal</td>
<td>overkill</td>
</tr>
<tr>
<td>Routing requests to the right model or agent</td>
<td>✅ ideal</td>
<td>✅ common today</td>
</tr>
<tr>
<td>Agent guardrails ("is this tool call dangerous?")</td>
<td>✅ ideal, but pair with deterministic checks</td>
<td>works, but slow</td>
</tr>
<tr>
<td>Fraud, spam, or moderation triage at scale</td>
<td>✅ ideal</td>
<td>too costly</td>
</tr>
<tr>
<td>Real-time in-app decisions (games, bidding)</td>
<td>✅ only option</td>
<td>impossible (3–30s latency)</td>
</tr>
<tr>
<td>Evaluating agent outputs (LLM-as-judge)</td>
<td>✅ promising</td>
<td>✅ common today</td>
</tr>
<tr>
<td>Writing an email, essay, or code</td>
<td>❌ can't generate text</td>
<td>✅</td>
</tr>
<tr>
<td>Counting, arithmetic, exact dates</td>
<td>❌ weak</td>
<td>✅ (or use code)</td>
</tr>
<tr>
<td>Novel problem, open-ended answer</td>
<td>❌ answers must be predefined</td>
<td>✅</td>
</tr>
</tbody></table>
<p>The punchline: the winning architecture is hybrid. System One routes, scores, and guards at high volume. System Two handles the escalated hard cases. Plain code does arithmetic.</p>
<p><strong>The constant trap:</strong> Archestra tested Jev on 100 real agent tool calls and found that 79% of calls were harmless and stayed local. A classifier with 75% accuracy is therefore <em>worse than a hardcoded "benign" constant</em>. Accuracy alone is meaningless — you need to know the base rate.</p>
<hr />
<h2>Getting Jev Into Your Stack</h2>
<p>There are several ways to call Jev, and the right one depends on whether you want to manage credentials yourself or let a gateway handle them.</p>
<table>
<thead>
<tr>
<th>Route</th>
<th>Model ID</th>
<th>Best for</th>
</tr>
</thead>
<tbody><tr>
<td><strong>TypeSafe direct</strong></td>
<td><code>typesafe/jev-1.13</code></td>
<td>Tightest control, newest features first</td>
</tr>
<tr>
<td><strong>OpenRouter</strong></td>
<td><code>typesafe/jev-1.13</code></td>
<td>One key for many models</td>
</tr>
<tr>
<td><strong>Vercel AI Gateway</strong></td>
<td><code>typesafe-ai/jev</code></td>
<td>Existing Vercel/Next.js apps, zero data retention</td>
</tr>
<tr>
<td><strong>LiteLLM proxy</strong></td>
<td><code>typesafe/jev-1.13</code></td>
<td>Enterprise routing, cost tracking, multi-provider</td>
</tr>
</tbody></table>
<h3>If you're already running a gateway</h3>
<p>This is where Jev's adoption story gets interesting, because gateways didn't just pass it through — they built on it.</p>
<p><strong>LiteLLM</strong> added Jev in <code>v1.103.0-rc</code> and proxies its <code>/systemone</code> endpoint with logging and cost tracking. Your TypeSafe key stays on the proxy; clients only need a LiteLLM virtual key. Spend is logged under the versioned model TypeSafe reports (<code>typesafe/jev-1.13.0</code>).</p>
<p>But LiteLLM went further and uses Jev for two jobs inside the proxy itself:</p>
<ul>
<li><p><strong>Context compaction.</strong> LiteLLM asks Jev whether older tool results are still relevant. If Jev scores a result below 0.2, it replaces the result with a short notice — saving input tokens on long agent runs.</p>
</li>
<li><p><strong>Auto Router classification.</strong> LiteLLM's router uses Jev to decide which model tier a request belongs to. In their benchmark, Jev matched expected tiers on 95% of calls versus 73.75% for Claude Haiku, at 5.43x the speed and 96% lower cost.</p>
</li>
</ul>
<p><strong>Vercel AI Gateway</strong> exposes Jev with zero data retention and no output token charge. Vercel's AI SDK ships an experimental <code>evaluate</code> API that surfaces typed answers directly in TypeScript — so a department choice can select a support queue, a severity score can influence priority, and an uncertain result can trigger human review.</p>
<p><strong>Cloudflare Workers AI</strong>, <strong>Netlify AI Gateway</strong>, and <strong>OpenRouter</strong> all serve Jev as well. Netlify's integration is zero-config: install <code>@typesafe-ai/sdk</code> in a Netlify Function, and AI Gateway handles credentials and billing.</p>
<p>The pattern across all of them is the same: Jev is cheap enough to call <em>inside</em> infrastructure, not just from application code. That's why routers use it as a classifier and proxies use it as a guardrail.</p>
<hr />
<h2>Limitations You Should Know</h2>
<p><strong>Typed ≠ correct.</strong> A wrong answer from your allowed list is still wrong.</p>
<p><strong>No explanations.</strong> That makes audits harder. Keep humans in the loop for high-stakes calls.</p>
<p><strong>Prompt injection is real and vendor-acknowledged.</strong> TypeSafe's own limitations page for Jev 1.13 states that "content written to adversarially steer the model, whether that is an injected instruction, a deliberately misleading framing, or text that argues for its own classification, can move the answer." An independent security benchmark found 10 false positives and 13 false negatives out of 662 prompt-injection tests. A guard built on Jev belongs <em>alongside</em> deterministic checks, not instead of them.</p>
<p><strong>Weak at arithmetic, dates, counting, and adversarial input.</strong> Leave those to code.</p>
<p><strong>Not a replacement for LLMs.</strong> It's a different tool for a different job.</p>
<hr />
<h2>The Takeaway</h2>
<p>System Two writes. System One decides. The future is both.</p>
<p>You're running 10 million classifications per day. A decision model gets 29 out of 30 right for a fraction of a cent each. A large LLM gets 30 out of 30 for a hundred times more. Where's your cutoff — and what would you never trust to a decision model?</p>
<hr />
<h2>References</h2>
<ul>
<li><p>TypeSafe AI — Jev API documentation and limitations page: <a href="https://docs.typesafe.ai/api">https://docs.typesafe.ai/api</a></p>
</li>
<li><p>LiteLLM — "TypeSafe Jev on LiteLLM" (Sep 20, 2026): <a href="https://docs.litellm.ai/blog/typesafe%5C_jev">https://docs.litellm.ai/blog/typesafe\_jev</a></p>
</li>
<li><p>LiteLLM — "Reduce agent context with TypeSafe Jev and LiteLLM" (Sep 18, 2026): <a href="https://docs.litellm.ai/blog/typesafe-jev-compaction">https://docs.litellm.ai/blog/typesafe-jev-compaction</a></p>
</li>
<li><p>LiteLLM — "JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost" (Sep 20, 2026): <a href="https://docs.litellm.ai/blog/jev-classifier">https://docs.litellm.ai/blog/jev-classifier</a></p>
</li>
<li><p>Vercel — "How to classify, route, and score with Jev and AI SDK": <a href="https://vercel.com/kb/guide/typesafe-jev-and-ai-sdk">https://vercel.com/kb/guide/typesafe-jev-and-ai-sdk</a></p>
</li>
<li><p>CryptoBriefing — "TypeSafe opens Jev AI to public after rapid adoption forces waitlist removal" (Sep 21, 2026): <a href="https://cryptobriefing.com/typesafe-jev-ai-public-access/">https://cryptobriefing.com/typesafe-jev-ai-public-access/</a></p>
</li>
<li><p>explainx.ai — "Is Jev's 200x-Faster, 400x-Cheaper Claim Actually True?" (Sep 19, 2026): <a href="https://www.explainx.ai/blog/jev-speed-cost-claims-fact-check-2026">https://www.explainx.ai/blog/jev-speed-cost-claims-fact-check-2026</a></p>
</li>
<li><p>Archestra — "We Tested Jev on 100 Real Agent Calls. How Easy Is It To Beat a Constant?" (Sep 21, 2026): <a href="https://archestra.ai">https://archestra.ai</a></p>
</li>
<li><p>VentureBeat — "Jev AI agent security: Prompt injection risk" (Sep 21, 2026): <a href="https://venturebeat.com">https://venturebeat.com</a></p>
</li>
<li><p>Layer3Labs — "Jev Benchmarks: How Accurate Is TypeSafe AI's Model?" (Sep 22, 2026): <a href="https://www.layer3labs.io">https://www.layer3labs.io</a></p>
</li>
<li><p>Daniel Kahneman — <em>Thinking, Fast and Slow</em> (2011)</p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[Memory for AI Agents: Architectures, Tools, and How to Choose the Right One]]></title><description><![CDATA[Imagine talking to someone who forgets everything the moment you stop speaking. Every conversation starts from zero. No context. No learning. No continuity.
That’s how most AI chatbots work today. The]]></description><link>https://nightthoughts.me/memory-for-ai-agents-architectures-tools-and-how-to-choose-the-right-one</link><guid isPermaLink="true">https://nightthoughts.me/memory-for-ai-agents-architectures-tools-and-how-to-choose-the-right-one</guid><category><![CDATA[AI]]></category><category><![CDATA[agentic AI]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[llm]]></category><category><![CDATA[ai memory]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 20 Sep 2026 13:53:16 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/caa6044f-101f-4a39-953b-8d0cbd6a1648.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine talking to someone who forgets everything the moment you stop speaking. Every conversation starts from zero. No context. No learning. No continuity.</p>
<p>That’s how most AI chatbots work today. They’re <strong>stateless</strong>—they have no memory of past interactions. But “agentic AI” is changing that. An <strong>agent</strong> is an AI that can take actions, use tools, and pursue goals over time. To do that well, it needs memory—just like a human.</p>
<p>This piece breaks down the memory problem for AI agents, the main architectural patterns, the tools available, and the criteria you should use to choose the right one. At the end, I’ll share a recommendation that stands out for production systems.</p>
<hr />
<h2>1. Why AI Agents Need Memory</h2>
<p>Without memory, an AI agent:</p>
<ul>
<li>Repeats mistakes.</li>
<li>Forgets user preferences.</li>
<li>Loses context across sessions.</li>
<li>Can’t learn from its own actions.</li>
<li>Feels like a stranger every time you talk to it.</li>
</ul>
<p>Memory turns a chatbot into a <strong>colleague</strong>. It lets the agent remember what was said, what was done, and what worked. That’s the foundation of agentic AI.</p>
<hr />
<h2>2. The Four Types of AI Memory (In Simple Terms)</h2>
<p>Most frameworks borrow from human cognition and split memory into four types:</p>
<table>
<thead>
<tr>
<th>Type</th>
<th>What It Holds</th>
<th>Example</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Working Memory</strong></td>
<td>The “right now”</td>
<td>Current conversation, active task</td>
</tr>
<tr>
<td><strong>Episodic Memory</strong></td>
<td>“What happened”</td>
<td>Past conversations, actions, outcomes</td>
</tr>
<tr>
<td><strong>Semantic Memory</strong></td>
<td>“Facts and knowledge”</td>
<td>User preferences, world facts</td>
</tr>
<tr>
<td><strong>Procedural Memory</strong></td>
<td>“How to do things”</td>
<td>Learned skills, successful action sequences</td>
</tr>
</tbody></table>
<p>A good agent memory system needs to handle all four—or at least the ones relevant to its job.</p>
<hr />
<h2>3. Architectural Patterns for Agent Memory</h2>
<p>There’s no single way to build memory. Here are the main patterns, from simplest to most advanced:</p>
<h3>a) Simple In-Process Memory</h3>
<p>Everything stays in the current conversation window. Easy, but the agent forgets once the session ends.</p>
<h3>b) External Vector Store</h3>
<p>The agent saves memories in a database and retrieves relevant ones using semantic search. Like a searchable notebook.</p>
<h3>c) Tiered Memory Architecture</h3>
<p>Layered storage: “hot” (fast, recent), “warm” (less frequent), and “cold” (archival). Balances speed and capacity.</p>
<h3>d) Knowledge Graph Memory</h3>
<p>Memories are stored as connected nodes (entities and relationships). Helps the agent reason about how facts relate.</p>
<p>Each pattern has trade-offs in cost, latency, complexity, and accuracy.</p>
<hr />
<h2>4. The Tool Landscape: A Quick List</h2>
<p>Here are the most talked-about memory tools for AI agents today:</p>
<ul>
<li><strong>Mem0</strong> – Plug-and-play conversational memory. Good for simple use cases.</li>
<li><strong>Letta (formerly MemGPT)</strong> – Pioneered agents that edit their own memory.</li>
<li><strong>Zep</strong> – Focuses on temporal reasoning—how memories change over time.</li>
<li><strong>LangChain / LangMem</strong> – Memory components inside a broader agent ecosystem.</li>
<li><strong>Memori</strong> – A SQL-native, structured memory layer that remembers actions, not just chat.</li>
</ul>
<p>Each has strengths. The right choice depends on your specific requirements.</p>
<hr />
<h2>5. How to Choose a Memory Tool</h2>
<p>Before picking a tool, ask:</p>
<ul>
<li><strong>What does it remember?</strong> Just chat, or also tool calls and outcomes?</li>
<li><strong>Where is memory stored?</strong> Proprietary vector store, or your own database?</li>
<li><strong>Can you audit it?</strong> Can you inspect and query memories directly?</li>
<li><strong>How much does it cost?</strong> Tokens per query matter at scale.</li>
<li><strong>How fast is it?</strong> Does memory creation slow down the agent?</li>
<li><strong>Can you deploy it your way?</strong> Cloud, on-prem, or bring your own database?</li>
<li><strong>Does it forget intelligently?</strong> Stale memories should fade; important ones should stay.</li>
<li><strong>Does it work with your inference stack?</strong> If you run local models, does the tool support them natively?</li>
</ul>
<p>These criteria will guide you to the right tool for your use case.</p>
<hr />
<h2>6. Evaluating the Options: Why Memori Stands Out</h2>
<p>After applying the criteria above, one tool consistently rises to the top for production agents: <strong>Memori</strong>.</p>
<p>Memori isn’t just another memory tool. It treats memory as a <strong>data structuring problem</strong>, not a text-stuffing problem. Here’s why that matters:</p>
<h3>✅ SQL-Native and Auditable</h3>
<p>Memori uses everyday databases like SQLite, PostgreSQL, or MongoDB. You can inspect, query, and audit memories directly. No black-box vector store.</p>
<h3>✅ Remembers Actions, Not Just Chat</h3>
<p>Most tools capture conversation. Memori also captures <strong>tool calls, execution paths, decisions, and outcomes</strong>. The agent learns from what it <em>does</em>, not only what it <em>says</em>.</p>
<h3>✅ Structured Memory</h3>
<p>It converts messy dialogue into clean, queryable facts and relationships. That makes recall more precise and less token-heavy.</p>
<h3>✅ Intelligent Recall + Decay Scoring</h3>
<p>Memories that are accessed often and recently get prioritized. Stale ones fade away. This keeps context relevant without manual cleanup.</p>
<h3>✅ Performance and Cost</h3>
<p>On the LoCoMo benchmark (a standard long-context memory test), Memori reportedly achieved <strong>81.95% accuracy</strong>—outperforming Mem0 (<del>62%), LangMem (</del>78%), and Zep (<del>79%). And it did this using only **</del>1,294 tokens per query**, roughly <strong>5% of the cost</strong> of stuffing the full conversation into the prompt.</p>
<h3>✅ Deployment Flexibility</h3>
<p>Use Memori Cloud, or bring your own database (BYODB) for full control. Great for enterprise, on-prem, or VPC deployments.</p>
<h3>✅ Native Ollama Integration — A Platform Capability, Not an Afterthought</h3>
<p>Memori is <strong>LLM-agnostic by design</strong>, and local OpenAI-compatible endpoints like <strong>Ollama, vLLM, and llama.cpp</strong> are treated as first-class providers. This matters because a growing share of agent development happens on local or self-hosted inference, yet many memory layers assume a cloud LLM.</p>
<p>Memori supports Ollama through three concrete mechanisms:</p>
<ol>
<li><strong>Universal provider support.</strong> Memori ships with tested examples for OpenAI, Azure OpenAI, LiteLLM, and Ollama. Its “any OpenAI-compatible” configuration path means most local inference servers work out of the box.</li>
<li><strong>LiteLLM as a bridge.</strong> Ollama is supported natively via LiteLLM, so the same memory pipeline that talks to GPT-4 or Claude can also talk to a locally running Llama or Qwen model without code changes.</li>
<li><strong>Local embeddings stay local.</strong> Memori’s retrieval layer supports Ollama embeddings (e.g., <code>nomic-embed-text</code>), so the entire memory pipeline—extraction, augmentation, and recall—can run entirely offline.</li>
</ol>
<p>For anyone building a local-first agent stack, this is the difference between a memory layer that <em>can</em> work with your inference engine and one that <em>does</em> work with it by design.</p>
<hr />
<h2>7. Multi-Agent Memory: Shared, Not Distributed</h2>
<p>When you have more than one AI agent, a key question is: <strong>what should they share, and what should stay private?</strong></p>
<p>Memori uses a <strong>shared memory database</strong> with clear rules about what each agent can see. Think of it like a shared notebook for all your agents, but with different sections that are locked or open.</p>
<p>There are three main levels of separation:</p>
<h3>👤 User Level (Shared Across All Agents)</h3>
<p>All agents that serve the same user share basic facts and preferences. For example, if you tell your <strong>support bot</strong> that you use PostgreSQL, your <strong>sales bot</strong> can later say, “I see you use PostgreSQL—here’s a product that works well with it.” The sales bot doesn’t need to ask you again.</p>
<p><strong>What’s shared at this level:</strong></p>
<ul>
<li>Facts (e.g., “Alice uses PostgreSQL”)</li>
<li>Preferences (e.g., “Alice prefers email over phone”)</li>
<li>Skills (e.g., “Alice knows Python”)</li>
<li>Knowledge graph (connections between facts)</li>
</ul>
<h3>🤖 Agent Level (Private to Each Agent)</h3>
<p>Each agent has its own private notes about how it works. The support bot and the sales bot each have their own attributes and conversation histories. The sales bot doesn’t see the full chat log from the support bot.</p>
<p><strong>What’s private at this level:</strong></p>
<ul>
<li>Attributes (e.g., the support bot’s tone settings)</li>
<li>Conversation history (each agent keeps its own record of what was said)</li>
</ul>
<h3>💬 Session Level (Private to Each Conversation)</h3>
<p>Even within the same agent, each conversation session is kept separate. If you start a new chat with the support bot, it remembers you use PostgreSQL (from the shared user memory), but it doesn’t remember the exact words from your previous chat session.</p>
<p><strong>What’s private at this level:</strong></p>
<ul>
<li>The specific messages exchanged in that session</li>
<li>The context of that particular conversation</li>
</ul>
<h3>A Simple Example</h3>
<p>Imagine you have two agents: a <strong>personal assistant</strong> and a <strong>shopping helper</strong>.</p>
<ol>
<li>You tell your personal assistant: “I love dark roast coffee.”</li>
<li>Later, you ask your shopping helper: “What coffee should I buy?”</li>
<li>The shopping helper can say: “Since you love dark roast, here are some options.”</li>
<li>But the shopping helper cannot see the entire conversation you had with the personal assistant—only the fact that you love dark roast.</li>
</ol>
<p>That’s the balance: <strong>shared knowledge, private conversations.</strong></p>
<h3>Summary Table</h3>
<table>
<thead>
<tr>
<th>What is it?</th>
<th>Who can see it?</th>
<th>Example</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Facts &amp; Preferences</strong></td>
<td>All agents for the same user</td>
<td>“Alice uses PostgreSQL”</td>
</tr>
<tr>
<td><strong>Skills</strong></td>
<td>All agents for the same user</td>
<td>“Alice knows Python”</td>
</tr>
<tr>
<td><strong>Attributes</strong></td>
<td>Only the specific agent</td>
<td>Support bot’s tone settings</td>
</tr>
<tr>
<td><strong>Conversation history</strong></td>
<td>Only the specific agent + session</td>
<td>The exact messages in one chat</td>
</tr>
</tbody></table>
<p>So Memori is not a peer-to-peer system where each agent has its own memory that syncs with others. It’s a <strong>central shared memory</strong> with smart rules about who sees what. That makes it easier to build multi-agent systems where agents collaborate without stepping on each other’s toes.</p>
<hr />
<h2>8. When Memori Might Not Be the Best Fit</h2>
<p>Memori is powerful, but it’s not for everyone. If you need a quick, plug-and-play memory for a simple chatbot, Mem0 might be easier. If you want an agent that rewrites its own memory in a research setting, Letta is interesting. But for <strong>production agents that need structured, auditable, cost-efficient memory across many sessions and multiple agents</strong>, Memori is a very strong candidate.</p>
<hr />
<h2>9. The Big Takeaway</h2>
<p>Memory is not just a feature. It’s a <strong>fundamental architecture problem</strong>.</p>
<p>For AI agents to be truly useful over long periods, they need a well-designed memory system that can store, retrieve, and forget information intelligently. Stop treating memory as an afterthought. Design it as a first-class part of your agentic AI system.</p>
<p>And if you want a memory layer that’s structured, SQL-native, action-aware, cost-efficient, <strong>Ollama-compatible</strong>, and <strong>shared across multiple agents with smart isolation</strong>—<strong>Memori</strong> is a very strong place to start.</p>
<hr />
<h2>References</h2>
<ul>
<li><p><strong>Cognitive architecture for AI memory.</strong> “Your Model Has Humanity’s Cortex. It Needs Its Own Hippocampus.” Medium. <a href="https://medium.com/codetodeploy/your-model-has-humanitys-cortex-it-needs-its-own-hippocampus-36a4b9676c8d">https://medium.com/codetodeploy/your-model-has-humanitys-cortex-it-needs-its-own-hippocampus-36a4b9676c8d</a></p>
</li>
<li><p><strong>Local AI workstation setup.</strong> “I Turned My Mac Mini Into a Local AI Workstation — Here’s Exactly How.” Hashnode. <a href="https://nightthoughts.hashnode.dev/i-turned-my-mac-mini-into-a-local-ai-workstation-here-s-exactly-how">https://nightthoughts.hashnode.dev/i-turned-my-mac-mini-into-a-local-ai-workstation-here-s-exactly-how</a></p>
</li>
<li><p><strong>Memori overview.</strong> Memori Labs. “Introducing Memori Cloud: Fully Hosted SQL-Native Memory Layer for AI Agents.” <a href="https://memorilabs.ai/blog/launching-memori-cloud/">https://memorilabs.ai/blog/launching-memori-cloud/</a></p>
</li>
<li><p><strong>Memori multi-user isolation model.</strong> Memori Docs. “Multi-User Support.” <a href="https://memorilabs.ai/docs/memori-cloud/concepts/multi-user-support/">https://memorilabs.ai/docs/memori-cloud/concepts/multi-user-support/</a></p>
</li>
<li><p><strong>Memori LoCoMo benchmark results.</strong> Memori Labs. “Releasing Our LoCoMo Benchmark Paper.” <a href="https://memorilabs.ai/blog/memori-locomo-paper-results/">https://memorilabs.ai/blog/memori-locomo-paper-results/</a></p>
</li>
<li><p><strong>Memori LLM provider support (Ollama, vLLM, llama.cpp).</strong> Memori Manual — Doramagic. <a href="https://doramagic.ai/en/projects/memori/manual/">https://doramagic.ai/en/projects/memori/manual/</a></p>
</li>
<li><p><strong>Memory tool comparison (Letta, Mem0, Zep, LangMem).</strong> LoCoMo Benchmark Results. Hugging Face. <a href="https://huggingface.co/datasets/rovemark/locomo-benchmark-results/blob/main/README.md">https://huggingface.co/datasets/rovemark/locomo-benchmark-results/blob/main/README.md</a></p>
</li>
</ul>
<hr />
<p><em>Note: Benchmark numbers are based on Memori’s published results. Always verify with your own use case before choosing a tool.</em></p>
]]></content:encoded></item><item><title><![CDATA[What Is LLM Routing? A Simple Guide to Saving Money on AI]]></title><description><![CDATA[Most teams pick one AI model and use it for everything.
A short "summarize this" request costs the same as a complex "debug this code" request. That's wasteful — but the waste is invisible because you]]></description><link>https://nightthoughts.me/what-is-llm-routing-a-simple-guide-to-saving-money-on-ai</link><guid isPermaLink="true">https://nightthoughts.me/what-is-llm-routing-a-simple-guide-to-saving-money-on-ai</guid><category><![CDATA[cost-optimisation]]></category><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[local ai]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sun, 20 Sep 2026 03:46:20 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/a3396002-de6f-4cb1-b5cf-204309e3767c.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most teams pick one AI model and use it for everything.</p>
<p>A short "summarize this" request costs the same as a complex "debug this code" request. That's wasteful — but the waste is invisible because your bill doesn't tell you which requests were easy.</p>
<p>Think of it like taking a taxi for every trip. Even a trip to the corner store. A router is like choosing between a bike, a bus, or a taxi depending on the trip.</p>
<p>This guide explains what routing is, why it matters, and what tools exist. No code, no math. Just plain ideas.</p>
<hr />
<h2>What Is a Gateway? What Is a Router?</h2>
<p>These two words get mixed up all the time. Here's the difference.</p>
<p><strong>Gateway</strong> = the front door. It handles logins, API keys, logging, and sending your request to the right place.</p>
<p><strong>Router</strong> = the decision-maker inside the front door. It looks at your request and decides which model should handle it.</p>
<p><strong>Failover</strong> = if the chosen model is down or slow, the gateway tries another one automatically.</p>
<p>A router is like a GPS that reroutes you when there's traffic. It's not just picking the fastest route — it's picking a route that still works when the main road is closed.</p>
<p>Most gateways include a router. Some routers work on their own.</p>
<hr />
<h2>The Hidden Risk of One Model</h2>
<p>Routing isn't just about cost. It's about staying online.</p>
<p>In 2026, every major API provider has had outages. OpenAI, Anthropic, Google, and AWS have all gone down for hours. If your app depends on one model, your app goes down when that model does.</p>
<p>A local model on your own hardware is not immune either. Your machine can crash, overheat, or lose power.</p>
<p>The fix is the same as the cost fix: have more than one model, and a layer that can switch between them.</p>
<p>Teams that added routing for cost reasons discovered it also saved them during outages. Teams that added it for resiliency discovered it also cut their bill. Either reason is enough.</p>
<hr />
<h2>The Tricky Part: Context Sharing</h2>
<p>Routing sounds simple. Send easy requests to cheap models, hard requests to expensive ones.</p>
<p>But there's a catch most beginners don't see coming: <strong>models don't share memory.</strong></p>
<p>Imagine switching drivers in the middle of a race. The new driver gets in the car and asks, "Where are we going? What happened so far?" If nobody tells them, they're lost.</p>
<p>That's what happens when you switch models mid-conversation.</p>
<p><strong>Every model has a limit on how much it can remember at once.</strong> This is called the context window. Some models remember a lot. Some remember very little. If your conversation is too long for the new model, it either fails or forgets the beginning.</p>
<p><strong>Here's a real example.</strong> You ask your local model to "fix this function and run the tests." It calls a tool. The tool result comes back. Then you switch to a cloud model for the next step. The cloud model has no idea what function you're talking about, what tests ran, or what failed.</p>
<p><strong>Three simple rules to avoid this:</strong></p>
<ol>
<li><strong>Don't switch models in the middle of a task.</strong> Finish the task, then switch if needed.</li>
<li><strong>Keep conversations short when switching.</strong> Long chats lose more context.</li>
<li><strong>If you must switch, pass a short summary.</strong> Even one sentence like "We're fixing a Python function that sorts a list" helps a lot.</li>
</ol>
<p>There's also a money angle. When you stay on the same model, providers give you a discount on repeated text. Switch models, and you lose that discount. So switching isn't just a context problem — it can cost more too.</p>
<p>For a personal setup like a Mac Mini, you don't need to solve all of this. Just know it exists. The simplest fix is: <strong>pick one model per task, and stick with it until the task is done.</strong></p>
<hr />
<h2>The Four Approaches</h2>
<p>Each approach has a cost benefit and a resiliency benefit. Cascade gives you both in one design.</p>
<p><strong>Cascade</strong> — Try a cheap model first. If it fails or the answer seems weak, try an expensive one. Simple to set up. The expensive model is already your fallback. Best for beginners.</p>
<p><strong>Semantic</strong> — Guess how hard the request is before sending it. Easy requests go to cheap models, hard ones go to expensive models. More accurate than cascade, but needs a classifier. You configure failover separately.</p>
<p><strong>Cost-Aware</strong> — Predict how much the request will cost before sending it. Pick the cheapest model that can handle it. Best for budget control. But pure cost focus means no built-in fallback.</p>
<p><strong>Learned</strong> — Train a small AI model to make routing decisions. Highest ceiling, but needs data and machine learning skills. Resiliency depends on whether your training data includes failures.</p>
<table>
<thead>
<tr>
<th>Approach</th>
<th>Cost</th>
<th>Resiliency</th>
</tr>
</thead>
<tbody><tr>
<td>Cascade</td>
<td>High</td>
<td>High — built in</td>
</tr>
<tr>
<td>Semantic</td>
<td>High</td>
<td>Medium — needs setup</td>
</tr>
<tr>
<td>Cost-Aware</td>
<td>Highest</td>
<td>Low — no fallback</td>
</tr>
<tr>
<td>Learned</td>
<td>High</td>
<td>Medium — depends</td>
</tr>
</tbody></table>
<hr />
<h2>The Tools — What's Actually Out There</h2>
<p>Here's what exists in 2026. Split into two categories: gateways and routers.</p>
<h3>Gateways</h3>
<table>
<thead>
<tr>
<th>Tool</th>
<th>Cost</th>
<th>Failover</th>
<th>Best For</th>
</tr>
</thead>
<tbody><tr>
<td>LiteLLM</td>
<td>Free OSS (self-hosted)</td>
<td>Yes</td>
<td>Control on a budget</td>
</tr>
<tr>
<td>OpenRouter</td>
<td>5.5% fee on top-ups</td>
<td>Yes</td>
<td>Fastest start</td>
</tr>
<tr>
<td>Portkey</td>
<td>Free 10K req/mo; Pro $49/mo</td>
<td>Yes</td>
<td>Governance</td>
</tr>
<tr>
<td>Cloudflare AI Gateway</td>
<td>Free tier; pay beyond</td>
<td>Yes</td>
<td>Apps on Cloudflare</td>
</tr>
<tr>
<td>Kong AI Gateway</td>
<td>Free OSS core; from ~$500/mo</td>
<td>Yes</td>
<td>Enterprises on Kong</td>
</tr>
</tbody></table>
<h3>Routers</h3>
<table>
<thead>
<tr>
<th>Tool</th>
<th>Cost</th>
<th>Failover</th>
<th>Best For</th>
</tr>
</thead>
<tbody><tr>
<td>RouteLLM</td>
<td>Free OSS</td>
<td>No</td>
<td>Simple cost routing</td>
</tr>
<tr>
<td>vLLM Semantic Router</td>
<td>Free OSS + GPU</td>
<td>Yes</td>
<td>High-volume serving</td>
</tr>
<tr>
<td>NVIDIA Switchyard</td>
<td>Free OSS</td>
<td>Yes</td>
<td>Agent workflows</td>
</tr>
</tbody></table>
<p>Most teams need a gateway with routing built in, not a standalone router. LiteLLM is the most common starting point because it's free, open source, and includes routing.</p>
<hr />
<h2>The Numbers — And Why They Vary</h2>
<p>Savings numbers are real but come from specific setups. Here's what the data actually shows.</p>
<table>
<thead>
<tr>
<th>Source</th>
<th>Savings</th>
<th>Setup</th>
<th>Shows</th>
</tr>
</thead>
<tbody><tr>
<td>LiteLLM production</td>
<td>51% → 60.7%</td>
<td>450+ users, 270k requests</td>
<td>Savings grow with tuning</td>
</tr>
<tr>
<td>LiteLLM production</td>
<td>95% served by non-flagship</td>
<td>Auto-Router classification</td>
<td>Most requests are easy</td>
</tr>
<tr>
<td>Subtask routing</td>
<td>46% cheaper</td>
<td>Coding agent</td>
<td>Routing by task type works</td>
</tr>
<tr>
<td>AT&amp;T production</td>
<td>Up to 90%</td>
<td>45B tokens/day</td>
<td>Scale amplifies savings</td>
</tr>
</tbody></table>
<p><strong>Why the numbers vary:</strong></p>
<ul>
<li><strong>Your workload matters.</strong> If most of your requests are complex, routing saves less.</li>
<li><strong>Your model pair matters.</strong> Routing between two cheap models saves little. Routing between cheap and expensive saves a lot.</li>
<li><strong>Tuning matters.</strong> The LiteLLM deployment started at 42.9% savings and climbed to 60.7% over four months as the team adjusted which requests went where.</li>
</ul>
<p>Routing is not magic. It works when you have a mix of easy and hard requests. If every request is hard, it won't help.</p>
<hr />
<h2>What's Next</h2>
<p>In the next piece, we'll add a routing layer to the Mac Mini setup. Not just to save money — but to keep working when the cloud goes down.</p>
<p>We'll install LiteLLM, connect local and cloud models, set up failover, and measure what it actually saves on a real workload.</p>
<hr />
<h2>References</h2>
<ol>
<li>LiteLLM. "51% Cost Savings Reported From a Live Production Deployment." <em>LiteLLM Blog</em>, August 10, 2026.</li>
<li>LiteLLM. "Subtask-Specific Routing: Same Quality, 46% Less Cost." <em>LiteLLM Blog</em>, September 7, 2026.</li>
<li>Contabo. "Best LLM Gateways in 2026: Top LiteLLM Alternatives." <em>Contabo Blog</em>, June 12, 2026.</li>
<li>OpenRouter. "Pricing." <em>OpenRouter</em>, 2026.</li>
<li>Portkey. "Feature Comparison." <em>Portkey Docs</em>, July 27, 2026.</li>
<li>Mobile World Live. "Analysis: AT&amp;T bets on routing, open model to tame AI costs." <em>Mobile World Live</em>, July 28, 2026.</li>
</ol>
]]></content:encoded></item><item><title><![CDATA[Open-Weight vs Proprietary AI: Cost, Licensing, and Accuracy Compared]]></title><description><![CDATA[Everyone is talking about open-weight AI models. You can download them. You can run them. You can even build products on them.
So they must be free, right?
Not even close.
This study breaks down what ]]></description><link>https://nightthoughts.me/open-weight-vs-proprietary-ai-cost-licensing-and-accuracy-compared</link><guid isPermaLink="true">https://nightthoughts.me/open-weight-vs-proprietary-ai-cost-licensing-and-accuracy-compared</guid><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[api]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Sat, 19 Sep 2026 16:04:22 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/ee991715-cda3-40a3-b41c-d442b2801dce.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everyone is talking about open-weight AI models. You can download them. You can run them. You can even build products on them.</p>
<p>So they must be free, right?</p>
<p>Not even close.</p>
<p>This study breaks down what AI models actually cost in 2026 — split into <strong>open-weight</strong> and <strong>proprietary</strong> models — and shows why "open" is not the same as "free."</p>
<p>But before we dive in, let's understand the basic unit of AI pricing: the token.</p>
<hr />
<h2>Before We Start: What Is a Token?</h2>
<p>AI models don't read words like humans do. They break text into small pieces called <strong>tokens</strong>.</p>
<p>A token is roughly <strong>4 characters</strong> or about <strong>¾ of a word</strong> in English.</p>
<table>
<thead>
<tr>
<th>Text</th>
<th>Tokens</th>
</tr>
</thead>
<tbody><tr>
<td>"Hello"</td>
<td>1</td>
</tr>
<tr>
<td>"Artificial intelligence"</td>
<td>2</td>
</tr>
<tr>
<td>"ChatGPT is amazing"</td>
<td>4</td>
</tr>
<tr>
<td>"I love learning about AI models"</td>
<td>7</td>
</tr>
</tbody></table>
<p>When you use an AI model through an API, you pay per <strong>1 million tokens</strong>. There are two types:</p>
<ul>
<li><p><strong>Input tokens</strong> — the text you send to the model</p>
</li>
<li><p><strong>Output tokens</strong> — the text the model sends back</p>
</li>
</ul>
<p>Output tokens almost always cost more than input tokens — usually <strong>5 to 6 times more</strong>. Generating new text takes more computing power than reading existing text.</p>
<p>So when you see "\(1.25 input / \)10 output per 1M tokens," you pay $1.25 for every million tokens you send, and $10 for every million the model generates. Long answers get expensive fast.</p>
<p>Now that you understand tokens, let's look at the real costs.</p>
<hr />
<h2>Part 1: Open-Weight Models Are Not Free</h2>
<p>Open-weight models are models whose parameters are publicly downloadable. Anyone can grab them.</p>
<p>But three hidden costs come with them:</p>
<ol>
<li><p><strong>Hardware.</strong> Big models need big machines.</p>
</li>
<li><p><strong>Licenses.</strong> Many "open" licenses have strings attached.</p>
</li>
<li><p><strong>People.</strong> Someone has to run and maintain the servers.</p>
</li>
</ol>
<h3>The License Trap</h3>
<table>
<thead>
<tr>
<th>Model</th>
<th>License</th>
<th>Commercial Use &amp; Catch</th>
</tr>
</thead>
<tbody><tr>
<td>DeepSeek R1 / V3</td>
<td>MIT</td>
<td>Free, no restrictions</td>
</tr>
<tr>
<td>Gemma 4</td>
<td>Apache 2.0</td>
<td>Free, no restrictions</td>
</tr>
<tr>
<td>Phi-4</td>
<td>MIT</td>
<td>Free, no restrictions</td>
</tr>
<tr>
<td>GLM-5.2</td>
<td>MIT</td>
<td>Free, no restrictions</td>
</tr>
<tr>
<td>OLMo</td>
<td>Apache 2.0</td>
<td>Free, no restrictions</td>
</tr>
<tr>
<td>Llama 4</td>
<td>Llama Community</td>
<td>Free until 700M monthly users</td>
</tr>
<tr>
<td>Mistral Medium 3.5</td>
<td>Modified MIT</td>
<td>Commercial supply restriction</td>
</tr>
<tr>
<td>Kimi K3</td>
<td>Custom</td>
<td>Revenue share up to 30% for MaaS</td>
</tr>
<tr>
<td>Grok-2</td>
<td>Custom</td>
<td>Free only under $1M/year revenue</td>
</tr>
<tr>
<td>StableLM</td>
<td>Custom</td>
<td>Free only under $1M/year revenue</td>
</tr>
<tr>
<td>Command R</td>
<td>CC-BY-NC</td>
<td>Research only. No commercial use</td>
</tr>
</tbody></table>
<p><strong>The lesson:</strong> MIT and Apache 2.0 licenses are genuinely free. Custom licenses are not. They look open, but they're really "free until you get big."</p>
<h3>The Hardware Bill</h3>
<table>
<thead>
<tr>
<th>Model</th>
<th>Size</th>
<th>Hardware Needed</th>
<th>Rough Cost</th>
</tr>
</thead>
<tbody><tr>
<td>Kimi K3</td>
<td>2.8T params</td>
<td>Two 8-GPU H200 servers</td>
<td>$320K–$420K+</td>
</tr>
<tr>
<td>Llama 4 Maverick</td>
<td>400B</td>
<td>8× H100 GPUs</td>
<td>$80K–$120K</td>
</tr>
<tr>
<td>DeepSeek V3</td>
<td>671B</td>
<td>4–8 datacenter GPUs</td>
<td>$50K–$80K</td>
</tr>
<tr>
<td>Gemma 4 31B</td>
<td>31B</td>
<td>One gaming GPU</td>
<td>~$2,000</td>
</tr>
<tr>
<td>Phi-4</td>
<td>14B</td>
<td>One gaming GPU</td>
<td>~$2,000</td>
</tr>
</tbody></table>
<p>Small models really can be nearly free. Frontier models cannot.</p>
<h3>The Break-Even Point</h3>
<p>Renting an 8-GPU NVIDIA H100 node on a specialized cloud provider costs roughly $18,000 per month, though major hyperscalers can charge over $70,000 for the same setup.</p>
<table>
<thead>
<tr>
<th>Daily Usage</th>
<th>API Cost/Month</th>
<th>Self-Host Cost/Month</th>
<th>Cheaper</th>
</tr>
</thead>
<tbody><tr>
<td>1M tokens</td>
<td>$150</td>
<td>$18,221</td>
<td>API</td>
</tr>
<tr>
<td>50M tokens</td>
<td>$7,500</td>
<td>$18,221</td>
<td>API</td>
</tr>
<tr>
<td>100M tokens</td>
<td>$15,000</td>
<td>$18,221</td>
<td>API</td>
</tr>
<tr>
<td>500M+ tokens</td>
<td>—</td>
<td>—</td>
<td>Self-host</td>
</tr>
</tbody></table>
<p><strong>Simple rule:</strong> Unless you're processing over <strong>100–500 million tokens per day</strong>, self-hosting a big open model costs more than just paying for an API.</p>
<hr />
<h2>Part 2: Proprietary Models — Pay As You Go</h2>
<p>Proprietary models are rented, not owned. You send text, you get answers, you pay per token.</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Company</th>
<th>Input $/1M</th>
<th>Output $/1M</th>
</tr>
</thead>
<tbody><tr>
<td>GPT-5.5</td>
<td>OpenAI</td>
<td>$5.00</td>
<td>$45.00</td>
</tr>
<tr>
<td>GPT-5</td>
<td>OpenAI</td>
<td>$1.25</td>
<td>$10.00</td>
</tr>
<tr>
<td>GPT-5 mini</td>
<td>OpenAI</td>
<td>$0.25</td>
<td>$2.00</td>
</tr>
<tr>
<td>GPT-5 nano</td>
<td>OpenAI</td>
<td>$0.05</td>
<td>$0.40</td>
</tr>
<tr>
<td>Claude Opus 5</td>
<td>Anthropic</td>
<td>$5.00</td>
<td>$25.00</td>
</tr>
<tr>
<td>Claude Sonnet 5</td>
<td>Anthropic</td>
<td>$2.00</td>
<td>$10.00</td>
</tr>
<tr>
<td>Gemini 3.8 Flash</td>
<td>Google</td>
<td>$0.75</td>
<td>$3.75</td>
</tr>
<tr>
<td>DeepSeek V4 Pro</td>
<td>DeepSeek</td>
<td>$0.66</td>
<td>$1.98</td>
</tr>
<tr>
<td>Kimi K3</td>
<td>Moonshot AI</td>
<td>$3.00</td>
<td>$15.00</td>
</tr>
</tbody></table>
<p>The good news: <strong>no license negotiations, no user caps, no hardware.</strong> You just pay the bill.</p>
<p>Two things to watch:</p>
<ul>
<li><p><strong>Output costs more than input</strong> — usually 5–6× more.</p>
</li>
<li><p><strong>Batch processing is 50% cheaper</strong> if you don't need answers immediately.</p>
</li>
</ul>
<hr />
<h2>Part 3: Cost vs. Accuracy</h2>
<p>Here's the big question: how much accuracy do you lose by going cheap?</p>
<h3>General Knowledge (MMLU-Pro)</h3>
<table>
<thead>
<tr>
<th>Model</th>
<th>Type</th>
<th>Score</th>
<th>Input Cost</th>
</tr>
</thead>
<tbody><tr>
<td>Qwen3.7 Max</td>
<td>Proprietary</td>
<td>89.6%</td>
<td>API only</td>
</tr>
<tr>
<td>Claude Opus 4.5</td>
<td>Proprietary</td>
<td>89.5%</td>
<td>$5.00</td>
</tr>
<tr>
<td>Qwen3.5 397B</td>
<td>Open weight</td>
<td>87.8%</td>
<td>Self-host</td>
</tr>
<tr>
<td>Kimi K2.5</td>
<td>Open weight</td>
<td>87.1%</td>
<td>$3.00</td>
</tr>
<tr>
<td>DeepSeek V4 Pro</td>
<td>Open weight</td>
<td>87.1%</td>
<td>$0.66</td>
</tr>
<tr>
<td>Gemma 4 31B</td>
<td>Open weight</td>
<td>~85%</td>
<td>Nearly free</td>
</tr>
</tbody></table>
<p>The gap is about <strong>2 points</strong> — but the price difference is <strong>40×</strong>.</p>
<h3>Real-World Coding (SWE-bench Verified)</h3>
<p>This is the hard one. It measures whether a model can actually fix real software bugs.</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Type</th>
<th>Score</th>
</tr>
</thead>
<tbody><tr>
<td>Claude Opus 5</td>
<td>Proprietary</td>
<td>96.0%</td>
</tr>
<tr>
<td>Claude Fable 5</td>
<td>Proprietary</td>
<td>95.0%</td>
</tr>
<tr>
<td>Ornith-1.5-397B</td>
<td>Open weight</td>
<td>86.0%</td>
</tr>
<tr>
<td>DeepSeek V4 Pro</td>
<td>Open weight</td>
<td>80.6%</td>
</tr>
<tr>
<td>MiniMax M3</td>
<td>Open weight</td>
<td>80.5%</td>
</tr>
<tr>
<td>Kimi K2.6</td>
<td>Open weight</td>
<td>~75%</td>
</tr>
</tbody></table>
<p>Here the gap is real. For genuinely hard work — complex coding, multi-step reasoning — proprietary models still win.</p>
<hr />
<h2>Part 4: So What Should You Do?</h2>
<p><strong>Use open weights when:</strong></p>
<ul>
<li><p>Your task is simple (summarizing, sorting, basic chat)</p>
</li>
<li><p>You need data to stay on your own servers</p>
</li>
<li><p>You're processing over 100M tokens/day</p>
</li>
<li><p>You want to avoid vendor lock-in</p>
</li>
</ul>
<p><strong>Use proprietary models when:</strong></p>
<ul>
<li><p>The task is hard (complex reasoning, agentic coding)</p>
</li>
<li><p>Your volume is low or medium</p>
</li>
<li><p>You need legal protection against IP claims</p>
</li>
<li><p>You don't want to hire engineers to run servers</p>
</li>
</ul>
<p><strong>The best answer for most teams: use both.</strong> Route easy tasks to cheap open models. Send hard tasks to frontier proprietary models.</p>
<hr />
<h2>The Bottom Line</h2>
<p>Open-weight models changed the game. They can cut your costs by <strong>80–95%</strong> while giving up only a few points of accuracy.</p>
<p>But three things stay true:</p>
<ol>
<li><p><strong>Downloading is free. Running is not.</strong> A frontier open model can cost $300K+ in hardware.</p>
</li>
<li><p><strong>"Open" licenses often have limits.</strong> Only MIT and Apache 2.0 are truly free.</p>
</li>
<li><p><strong>Someone has to maintain it.</strong> Budget for engineering time, not just GPUs.</p>
</li>
</ol>
<p>The smartest teams in 2026 aren't choosing sides. They're mixing both.</p>
<hr />
<h2>References</h2>
<ol>
<li><p>MillionMiner. "AI Server Price Guide 2026: Rent Vs Buy." September 9, 2026.</p>
</li>
<li><p>IntuitionLabs. "Open-Weight AI Model Licenses: Commercial Use Rules Explained." September 5, 2026.</p>
</li>
<li><p>Taskade. "10 Best Open-Source LLMs, August 2026." May 23, 2026.</p>
</li>
<li><p>BenchLM. "Self-Hosting vs API: Break-Even Calculator for Open-Weight LLMs." July 3, 2026.</p>
</li>
<li><p>OpenAI. "Pricing." OpenAI API, 2026.</p>
</li>
<li><p>Anthropic. "Claude Opus 5 Pricing." July 24, 2026.</p>
</li>
<li><p>Google. "What's new in Gemini 3.8 Flash." September 3, 2026.</p>
</li>
<li><p>BenchLM. "MMLU-Pro Leaderboard &amp; Scores — September 2026." September 9, 2026.</p>
</li>
<li><p>BenchLM. "SWE-bench Verified Leaderboard (September 2026)." September 10, 2026.</p>
</li>
<li><p>Meta. "Llama 4 Community License Agreement." April 5, 2025.</p>
</li>
<li><p>Hugging Face. "Kimi K3 License." July 30, 2026.</p>
</li>
<li><p>Reuters. "China's DeepSeek to make permanent 75% price cut on flagship V4-Pro AI model." May 23, 2026.</p>
</li>
</ol>
<p><em>All pricing and benchmark data reflects publicly available information as of September 2026. Verify current terms and pricing before making procurement decisions.</em></p>
]]></content:encoded></item><item><title><![CDATA[When AI Eats Itself: Understanding Model Collapse and AI Cannibalism]]></title><description><![CDATA[There is a snake in the AI industry, and it is eating its own tail.
For years, the recipe for building better AI seemed simple: scrape more data from the internet, train a bigger model, repeat. But we]]></description><link>https://nightthoughts.me/when-ai-eats-itself-understanding-model-collapse-and-ai-cannibalism</link><guid isPermaLink="true">https://nightthoughts.me/when-ai-eats-itself-understanding-model-collapse-and-ai-cannibalism</guid><category><![CDATA[AI]]></category><category><![CDATA[generative ai]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[Data Science]]></category><category><![CDATA[llm]]></category><dc:creator><![CDATA[Haithem Slimi]]></dc:creator><pubDate>Fri, 18 Sep 2026 21:30:57 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a92730f9a9aa7f72e74fdf4/514e6e4a-cb11-4a81-b31b-1f90124ecd8f.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There is a snake in the AI industry, and it is eating its own tail.</p>
<p>For years, the recipe for building better AI seemed simple: scrape more data from the internet, train a bigger model, repeat. But we have reached a strange turning point. A huge portion of the internet is now written by AI. And the next generation of AI is being trained on that very content. The models are eating themselves.</p>
<p>This is not a metaphor. It has a name, and it is a formally proven problem called <strong>model collapse</strong>.</p>
<h2>What the Researchers Found</h2>
<p>In July 2024, a team led by Ilia Shumailov at Oxford University, working with researchers from Cambridge and other institutions, published a landmark paper in the journal <em>Nature</em>. Their finding was stark and simple: if you train AI models on data generated by other AI models, they will break down in an irreversible way.</p>
<p>They called this effect <strong>"model collapse"</strong>.</p>
<p>The study showed that this does not just happen to one type of model. They observed the collapse in large language models (LLMs), in image models called variational autoencoders, and even in simple statistical models called Gaussian mixture models. The problem is baked into the very nature of how these systems learn.</p>
<h2>Why Does Collapse Happen?</h2>
<p>To understand why, you need to think about what a model actually learns. It is not learning "facts." It is learning <strong>statistical patterns</strong> from enormous piles of data. It learns what words tend to follow other words, what ideas tend to appear together.</p>
<p>A healthy dataset is full of variety. It contains common ideas that appear all the time. But it also contains rare things: minority languages, unusual turns of phrase, niche knowledge, strange edge cases. Researchers call these rare parts the <strong>"tails"</strong> of the data distribution.</p>
<p>Here is the problem: <strong>every model slightly overproduces the common stuff and underproduces the rare stuff</strong>. It is a small bias, but it exists in every generation of the model.</p>
<p>Now imagine you take that output — which is slightly skewed toward the average — and you train the <em>next</em> model on it. The rare content shrinks even more. You do this again, and again. The tails of the distribution disappear entirely.</p>
<p>What you are left with is a model that has forgotten the richness of the real world. It only knows the middle, the average, the bland. Its outputs become repetitive and eventually nonsensical.</p>
<h2>The Jackrabbit Problem</h2>
<p>The researchers gave a vivid demonstration of this. They took a small model called OPT-125m and fed it a passage of text about <strong>14th-century church towers</strong>.</p>
<p>Then they let it generate new text. They trained a new model on that output. Then another. And another. They did this for nine generations.</p>
<p>By generation nine, the model had completely forgotten about architecture. When asked about medieval church towers, it started listing species of <strong>jackrabbits</strong> — including fictional ones like "blue-tailed jackrabbits".</p>
<p>The model did not just get a fact wrong. It had lost the entire conceptual grounding of the original topic. It was hallucinating entirely.</p>
<h2>This Is Happening Right Now</h2>
<p>This is not a theoretical problem for the future. It is a problem for right now.</p>
<p>A study from the company Graphite analyzed over 55,000 English-language articles published between 2020 and 2026. They found that in the first quarter of 2026, an estimated <strong>49.9% of sampled articles were classified as mostly AI-generated</strong>.</p>
<p>In late 2024, that figure was around 48%. By late 2025, AI-generated articles briefly surpassed human-written ones.</p>
<p>Other studies paint a similar picture. One analysis of over 1.2 million webpages from late 2025 to early 2026 found that roughly <strong>35% contained wholly or partially AI-generated content</strong>. Estimates suggest that <strong>30–40% of the active web corpus is now synthetic</strong>.</p>
<p>This matters because those same models are trained on web-scraped data. They are increasingly ingesting content that other models wrote. The contamination is not a future scenario. It is the current state of the training pipeline.</p>
<h2>Can It Be Fixed?</h2>
<p>The research is not all doom. But the fixes are not simple.</p>
<p><strong>The most basic idea is to mix real data with synthetic data.</strong> A paper presented at an IEEE conference in late 2025 proposed a simple strategy: combine synthetic data and human-generated data so that the model does not degenerate. They derived mathematical conditions on the <em>ratio</em> of synthetic to human data needed to keep the system stable.</p>
<p>But there is a catch. The same paper notes that on a small GPT model, enriching synthetic data with just a <em>small</em> amount of human data <strong>may not be enough</strong> to prevent collapse. The balance matters.</p>
<p><strong>Another approach is to use detectors.</strong> Researchers at University College London trained a machine-generated text detector and proposed a method to up-sample likely human content in the training data. They found that this not only prevented model collapse but actually <strong>improved performance compared to training on purely human data</strong>.</p>
<p><strong>A third approach is to monitor entropy.</strong> As a model collapses, the entropy of its training data declines. The model stops generating novel content and starts memorizing its training samples. Researchers have proposed using this entropy decline as an indicator of degradation and selecting data to preserve diversity.</p>
<p><strong>A fourth idea is verification.</strong> Some researchers have shown that injecting information through an external verifier — whether a human or a better model — can prevent synthetic retraining from causing collapse.</p>
<p>The common thread in all these solutions is the same: <strong>human-generated data is irreplaceable</strong>. The original paper put it plainly: "the value of data collected about genuine human interactions with systems will be increasingly valuable in the presence of LLM-generated content in data crawled from the Internet".</p>
<h2>What This Means for You</h2>
<p>If you are building or training models, the lesson is straightforward: <strong>know where your data comes from</strong>. If you are scraping the web post-2022, you are training on synthetic content. Monitor your model's diversity and entropy during training.</p>
<p>If you are working with data pipelines, understand that genuine human data is becoming a strategic asset. The value of real conversations, original writing, and real-world observations is rising precisely because it is becoming scarce.</p>
<p>And if you are watching the industry broadly, the "free lunch" of scraping unlimited web data is ending. Sustainable AI development will require deliberate investment in data provenance, curation, and verification.</p>
<p>The snake can stop eating its tail. But it has to choose to do so before the tails disappear entirely.</p>
<hr />
<h2>Sources</h2>
<ul>
<li><p>Shumailov, I. et al. "AI models collapse when trained on recursively generated data." <em>Nature</em> 631, 755–759 (2024).</p>
</li>
<li><p>Graphite study on AI-generated articles, reported by Search Engine Land (2026).</p>
</li>
<li><p><a href="https://datasciencedojo.com/blog/ai-cannibalism-model-collapse/">Data Science Dojo, "AI Cannibalism: How Models Are Eating Themselves Into Collapse."</a></p>
</li>
<li><p>IEEE conference paper, "Preventing Model Collapse when Training LLMs with Synthetic Data" (2025).</p>
</li>
<li><p>UCL research on machine-generated text detection (2025).</p>
</li>
<li><p>NeurIPS paper, "A Closer Look at Model Collapse" (2025).</p>
</li>
</ul>
<p><em>Cover image source:</em> <a href="https://datasciencedojo.com/blog/ai-cannibalism-model-collapse/"><em>Data Science Dojo</em></a></p>
]]></content:encoded></item></channel></rss>