Ollama is a locally-installed service, written in Go, that downloads, caches and runs LLMs on your own machine and exposes them over a REST API on port 11434. LangChainGo, the Go port of LangChain, provides a low-level SDK that talks to that same REST API, so a Go program can drive a local model without hand-rolling HTTP requests.

Installing and running a model

Install Ollama using the instructions on the project's download page. Once the service is running, ollama run <modelname> starts an interactive session; the model is fetched and cached the first time it is requested. For the examples here, that means ollama run llama2.

Talking to the API directly

Because ollama runs in the background and serves HTTP, curl is enough to exercise it. Responses can take a while on machines without a strong GPU, so requests may set "stream": true to receive output incrementally instead of waiting for the full reply. A streamed response arrives as a series of JSON messages carrying "done": false; the final message sets "done": true.

The endpoint is not limited to text completion — the same REST interface can produce embeddings for a prompt.

Driving Ollama from Go

The Ollama README points to several programmatic options, the most common being LangChain and its ecosystem. LangChain offers high-level composition of LLM tasks plus per-model SDKs; LangChainGo is the Go equivalent.

A simple non-streaming completion goes through GenerateFromSinglePrompt. The same function supports streaming by accepting a streaming function as an option: that callback is invoked once per chunk as data arrives, and finally once with an empty chunk. The accumulated completion is still returned from GenerateFromSinglePrompt for callers that want it.

Embeddings are likewise available through the langchain package.

Reference implementation

The complete program accompanying this walkthrough is published on GitHub, and a followup post covers swapping in additional models such as Google's Gemma with the same setup.