<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Notes, Bhavuk Manocha</title><description>Notes on AI agents, local models and making them work at work.</description><link>https://bhavukmanocha.com/blog/</link><item><title>Build your own Instinct, or Muse, with Hermes Agent</title><link>https://bhavukmanocha.com/blog/build-your-own-instinct-with-hermes/</link><guid isPermaLink="true">https://bhavukmanocha.com/blog/build-your-own-instinct-with-hermes/</guid><description>Meta&apos;s Muse and the invite-only Instinct sell the same idea: an assistant that remembers you, lives in your chats and messages you first. Hermes Agent, an open-source agent from Nous Research, has every piece you need to build your own.</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Why plugging Claude straight into your data doesn&apos;t work</title><link>https://bhavukmanocha.com/blog/plugging-claude-into-your-warehouse/</link><guid isPermaLink="true">https://bhavukmanocha.com/blog/plugging-claude-into-your-warehouse/</guid><description>Connect a model to the warehouse, ask a question, get a confident number. Anthropic&apos;s own data team says that setup creates &quot;a false sense of precision&quot;. Their fix is mostly not about the model.</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate></item><item><title>A 35B model on an 8 GB graphics card</title><link>https://bhavukmanocha.com/blog/qwen36-35b-on-an-8gb-gpu/</link><guid isPermaLink="true">https://bhavukmanocha.com/blog/qwen36-35b-on-an-8gb-gpu/</guid><description>Qwen3.6-35B-A3B has 36 billion parameters and still runs on a mid-range GPU at around 40 tokens a second. The trick is knowing which 7% of the model has to be on the card, and parking the rest in system RAM.</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate></item></channel></rss>