Local Agentic AI Workflows with Hermes Agent and Ollama

Build a fully local, zero-cost agentic AI workflow using Hermes Agent and Ollama. Keep your files, code, and conversations on your own hardware.

Local Agentic AI Workflows with Hermes Agent and Ollama

In this article, you will learn how to build a fully local, zero-cost agentic AI workflow using Hermes Agent and Ollama, so that your files, code, and conversations never leave your own hardware.

Topics we will cover include:

  • How to install Ollama, choose the right local model for agentic work, and verify that the model is responding correctly before wiring anything else up.
  • How to configure Hermes Agent to use your local Ollama endpoint, and how to optimize context window size and model loading for real agentic tasks.
  • How to extend the setup with a Telegram gateway for remote access and a cloud fallback for questions the local model cannot handle well.

Local Agentic AI Workflows with Hermes + Ollama

A typical coding session against a cloud AI API runs somewhere between $0.60 and $0.80 depending on the provider, and a heavier session can climb to $5 to $20, according to Nous Research’s own cost breakdown for agentic work. That adds up fast for a hobbyist, a student, or anyone running frequent automation, and it comes with a second cost that is easy to overlook: every file, every question, every line of code gets sent to a third party’s servers.

This article builds the alternative: a genuinely local, zero-cost agentic AI workflow using Hermes Agent, an open-source AI agent from Nous Research, paired with Ollama for local model serving.

What Is Hermes Agent?

Hermes Agent is an open-source AI agent built by Nous Research, released under the MIT license and currently at version 0.21.1 as of this writing. It ships two ways: a native desktop app for macOS, Windows, and Linux, and a terminal-first CLI you install directly. What separates it from a basic chat interface is genuine agentic capability; it edits files, runs terminal commands, browses the web, and can delegate work to isolated sub-agents with their own conversations and tools.

A few features matter specifically for this article. Persistent memory means Hermes learns your projects over time and can auto-generate reusable skills from how it solved past problems, rather than starting from zero every session. Its messaging gateway connects the same agent and the same memory to Telegram, Discord, Slack, WhatsApp, and email. And its sandboxing system supports five different isolation backends — local, Docker, SSH, Singularity, and Modal — so commands it runs do not have to touch your host system directly if you would rather they did not.

What Is Ollama?

Ollama is the layer underneath Hermes in this setup: a tool that downloads, serves, and manages open-weight language models directly on your own hardware, exposing them through a local API that looks and behaves like a standard cloud LLM endpoint. That last detail matters more than it sounds: because Ollama’s API is OpenAI-compatible at /v1/chat/completions, Hermes can talk to a model running entirely on your laptop using the exact same integration path it would use for a cloud provider like OpenAI or Anthropic — just pointed at localhost instead of the internet.

The division of labor is clean: Ollama’s only job is running the model and answering requests for it. Hermes’ job is being the actual agent — deciding when to call a tool, editing a file, running a command, browsing the web, and interpreting what comes back. Neither one replaces the other, and this tutorial needs both.

What We’re Building

The concrete project for this article is a private, zero-cost local assistant that can organize and answer questions about a real folder of files on your machine, search the web when a question genuinely needs current information, and — once the core setup works — stay reachable from your phone via a Telegram bot when you are away from your desk. As a final layer, it will have a cloud fallback configured so genuinely hard questions still get answered well, while the other 90% of everyday use costs nothing and never leaves your machine.

Every section from here builds one real piece of that project, in the order you would actually build it.

What You Need

Hardware requirements scale with the model you plan to run, and it is worth knowing both ends of the range before choosing.

ComponentMinimumRecommended
RAM8 GB (for 3B models)32+ GB (for 27B+ models)
Storage5 GB free30+ GB (for multiple models)
CPU4 cores8+ cores
GPUNot requiredNVIDIA GPU with 8+ GB VRAM

CPU-only setups genuinely work; they are just slower. A 9B model on a modern 8-core CPU runs at roughly 10 tokens per second, while a 31B model on CPU drops to about 2 to 5 tokens per second, meaning each response can take 30 to 120 seconds. That is usable for a background assistant, less pleasant for an interactive back-and-forth, which is worth factoring into which model you pick.

Install Ollama and Pull a Model

Install Ollama with its official install script:

curl -fsSL https://ollama.com/install.sh | sh

Confirm it is actually running:

ollama --version
curl http://localhost:11434/api/tags   # Should return {"models":[]}

Expected output:

$ ollama --version
ollama version is 0.33.2
$ curl http://localhost:11434/api/tags
{"models":[]}

The first command checks that the binary is installed correctly. The second hits Ollama’s local API directly, and an empty models array is the expected, correct response at this point; it confirms the server is listening — you just have not downloaded a model into it yet.

Now pull a model. This is the single most consequential choice in the whole setup, because not every model can actually act as an agent:

ModelSize on DiskRAM NeededTool CallingBest For
gemma4:31b~20 GB24+ GBYesBest quality, strong tool use and reasoning
gemma2:27b~16 GB20+ GBNoConversational tasks, no tool use
gemma2:9b~5 GB8+ GBNoFast chat, Q&A, cannot call tools
llama3.2:3b~2 GB4+ GBNoLightweight quick answers only

That “Tool Calling” column is the whole ballgame for this project. Hermes is an agentic assistant specifically because it can call tools, edit a file, run a command, search the web, and a model without tool-call support can only chat back at you — it cannot actually take an action on your behalf, no matter how well it writes. For the file-organizing, web-searching assistant this article is building, that means gemma4:31b is the real starting point, not the smaller options.

ollama pull gemma4:31b

Once it is downloaded, confirm the model itself actually answers correctly:

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma4:31b",
    "messages": [{"role": "user", "content": "Say hello"}],
    "max_tokens": 50
  }'

A successful response returns a JSON object with a choices array containing the model’s reply, confirming that Ollama is serving the model correctly and that the OpenAI-compatible endpoint is ready for Hermes to connect to.