ollama-haskell
Industry-grade Haskell client for Ollama local LLMs
https://github.com/tusharad/ollama-haskell
| LTS Haskell 24.55: | 0.2.1.0@rev:2 |
| Stackage Nightly 2026-08-23: | 0.4.0.0 |
| Latest on Hackage: | 0.4.0.0 |
ollama-haskell-0.4.0.0@sha256:31f1f2244b2a8698055d679e6c3fa98678543c174749bf387487c93e8f876b6f,5876Module documentation for 0.4.0.0
ollama-haskell
Industry-grade, feature-complete, modern Haskell client library for the Ollama local LLM engine.
Features
- Client-Centric Architecture: Thread-safe
OllamaClienthandle with connection pooling and resource management (newClient,defaultClient,clientFromEnv,withClient). - First-Class Streaming:
conduit-based response streaming (chatStream,generateStream,pullStream,pushStream,createModelStream). - Model Context Protocol (MCP) Bridge: Bidirectional integration with
mcp-serverfor converting between Ollama tools and MCP tools, running MCP servers via stdio or HTTP (Ollama.MCP). - Generic JSON Schema Derivation: Automatically derive JSON schemas from Haskell data types via
GHC.GenericswithToSchemaandformatFor. - Complete API Surface: Text generation, chat completions, vector embeddings, model management (list, show, copy, delete, pull, push, create), and system endpoints.
- Structured Outputs DSL: Powerful
SchemaBuilderDSL (|+,|++,|!,|!!) for type-safe JSON Schema structured responses. - Function / Tool Calling: Full support for tool definitions (
Tool), tool calls (ToolCall), and execution results (toolResultMessage). - Thinking Models Support: Native support for reasoning models (
qwen3.5,deepseek-r1) withThink/ThinkingLeveltypes. - Environment & Auth Integration: Robust URL normalization for
OLLAMA_HOSTand bearer token support forOLLAMA_API_KEY. - Configurable Resilience: Flexible retry policies (
NoRetry,ConstantRetry,ExponentialRetry), custom timeouts, lifecycle callbacks, and structured logging. - Conversation Store: Transactional STM-backed
InMemoryStoreandConversationStoretypeclass for managing multi-turn chat sessions. - SDK Comparison Matrix: Detailed feature comparison against Python, JS/TS, and Go SDKs in doc/COMPARISON.md.
Installation
Add ollama-haskell to your .cabal file:
build-depends:
base >= 4.17 && < 5
, ollama-haskell >= 0.4.0.0
Or using Stack in package.yaml:
dependencies:
- ollama-haskell >= 0.4.0.0
Quick Start (5 Lines)
import Data.List.NonEmpty (NonEmpty ((:|)))
import Data.Text.IO qualified as TIO
import Ollama
main :: IO ()
main = do
client <- defaultClient
res <- chat client $ chatRequest "qwen3.5:2b" (userMessage "Why is the sky blue?" :| [])
case res of
Left err -> print err
Right resp -> mapM_ (TIO.putStrLn . messageContent) (crMessage resp)
Streaming Responses with Conduit
Stream LLM responses token-by-token as they generate:
import Data.List.NonEmpty (NonEmpty ((:|)))
import Data.Text.IO qualified as TIO
import Ollama
main :: IO ()
main = do
client <- defaultClient
let req = chatRequest "qwen3.5:2b" (userMessage "Count from 1 to 5." :| [])
-- Stream chunks directly into stdout or collect them
chunks <- collectStream (chatStream client req)
mapM_ (TIO.putStr . maybe "" messageContent . crMessage) chunks
putStrLn ""
Function & Tool Calling
Define function signatures and let the LLM execute structured tool calls:
import Data.List.NonEmpty (NonEmpty ((:|)))
import Ollama
calculatorTool :: Tool
calculatorTool = Tool "function" $ FunctionDef
{ fnName = "add"
, fnDescription = Just "Add two numbers"
, fnParameters = Just (FunctionParameters "object" Nothing (Just ["a", "b"]) Nothing Nothing Nothing)
, fnStrict = Just True
}
main :: IO ()
main = do
client <- defaultClient
let req = (chatRequest "qwen3.5:2b" (userMessage "What is 40 + 2?" :| []))
{ chatTools = Just [calculatorTool] }
res <- chat client req
case res of
Left err -> print err
Right resp -> print (crMessage resp)
Structured Outputs (JSON Schema DSL)
Enforce structured JSON output formats using SchemaBuilder:
import Data.Text.IO qualified as TIO
import Ollama
import Ollama.Types.Format.SchemaBuilder
personSchema :: Schema
personSchema = buildSchema $ emptyObject
|+ ("name", JString)
|+ ("age", JInteger)
|! "name"
main :: IO ()
main = do
client <- defaultClient
let req = (generateRequest "qwen3.5:2b" "Generate a person profile.")
{ genFormat = Just (SchemaFormat personSchema) }
res <- generate client req
case res of
Left err -> print err
Right resp -> TIO.putStrLn (grResponse resp)
Environment Variables & Configuration
Construct a client using environment variables (OLLAMA_HOST, OLLAMA_API_KEY):
main :: IO ()
main = do
client <- clientFromEnv
-- Automatically connects to OLLAMA_HOST with optional Authorization: Bearer header
...
Or configure custom retry policies and loggers:
customConfig :: OllamaClientConfig
customConfig = defaultConfig
{ configBaseUrl = "http://my-ollama-server:11434"
, configTimeout = 120
, configRetry = ExponentialRetry 3 1 -- 3 retries with exponential backoff
, configLogger = Just (\level msg -> putStrLn $ "[" <> show level <> "] " <> show msg)
}
main :: IO ()
main = withClient customConfig $ \client -> do
...
Documentation & SDK Comparison
- doc/COMPARISON.md — SDK Feature Matrix comparing
ollama-haskellwith Python, JS/TS, and Go SDKs. - CONTRIBUTING.md — Development setup, testing guidelines, and code style.
- CHANGELOG.md — Release notes and changelog.
- Hackage Documentation — Full Haddock reference.
License
MIT © 2024–2026 Tushar Adhatrao
Changes
Changelog
All notable changes to ollama-haskell will be documented in this file.
The format is based on Keep a Changelog,
and this project adheres to PVP (Haskell Package Versioning Policy).
[0.4.0.0] - 2026-08-20
Added
- Model Context Protocol (MCP) Integration (
Ollama.MCP):- Full bidirectional integration with the Hackage
mcp-serverpackage (mcp-server >= 0.2 && < 0.3). - Seamless conversion between Ollama function calling definitions (
Tool,ToolCall) and MCP definitions (ToolDefinition,ArgumentDefinition,Content,McpSchema). - Bridge functions:
toolToMcpDefinition,mcpDefinitionToTool,toolCallToMcpArgs,mcpContentToToolOutput. - Re-exported MCP server runners (
runMcpServerStdio,runMcpServerHttp,runMcpServerHttpWithConfig). - Dedicated unit test suite in
Test.Ollama.Unit.MCP.
- Full bidirectional integration with the Hackage
- Automatic JSON Schema Derivation (
Ollama.Types.Format.SchemaDerive):- Typeclasses
ToSchemaandToJsonTypeenabling generic derivation of JSON schemas directly from Haskell record types viaGHC.Generics. - Smart handling of optional fields (
Maybe aomitted fromrequired), nested records (JObject), lists (JArray), and simple sum enums (stringenum). - Helper functions
schemaForandformatForfor effortless integration withchat/generatestructured outputs. - Dedicated unit test suite in
Test.Ollama.Unit.SchemaDerive.
- Typeclasses
- Configurable Client Timeout:
- Support for custom request timeout intervals in
OllamaClientConfig(configTimeout).
- Support for custom request timeout intervals in
Changed
- PVP Compliance & Upper Bounds:
- Added strict upper bounds for
network-uri(>= 2.6 && < 2.8) andmcp-server(>= 0.2 && < 0.3). - Upgraded Stack resolvers and snapshot dependencies (
lts-21.25,lts-22.44,lts-23.28,lts-24.52,nightly).
- Added strict upper bounds for
[0.3.0.0] - 2026-08-04
Added
OllamaClientCore: Thread-safe client handle with automatic connection manager lifecycle management (newClient,defaultClient,clientFromEnv,withClient).- First-Class Streaming Pipeline:
conduit-based response streaming (chatStream,generateStream,pullStream,pushStream,createModelStream). - Stream Combinators:
collectStreamandfoldStreaminOllama.Streaming. - Structured Output DSL: Type-safe
SchemaBuilderDSL inOllama.Types.Format.SchemaBuilder(|+,|++,|!,|!!) for constructing JSON Schemas. - Thinking / Reasoning Models Support:
ThinkADT (ThinkEnabled,ThinkDisabled,ThinkLevel) andThinkingLevel(ThinkLow,ThinkMedium,ThinkHigh,ThinkMax) supporting models such asqwen3.5anddeepseek-r1. - Function / Tool Calling:
Tool,FunctionDef,FunctionParameters,ToolCall, andtoolResultMessagehelper. - Environment Resolution: Automatic
OLLAMA_HOSTparsing and normalization inclientFromEnvsupportinghost:port,http://host:port, and barehost. - Authorization & Headers: Support for
OLLAMA_API_KEYbearer tokens and customconfigHeaders. - Configurable Resilience:
RetryPolicyADT (NoRetry,ConstantRetry,ExponentialRetry), lifecycle callbacks (configOnStart,configOnSuccess,configOnError), and structured loggerconfigLogger. - Token Throughput Metrics: Metrics helpers
chatEvalTokensPerSecond,chatPromptEvalTokensPerSecond,evalTokensPerSecond,promptEvalTokensPerSecond,tokensPerSecond. - Testing Infrastructure: Built-in mock testing module
Ollama.Testing(newMockClient,withMockClient,mockGenerateResponse,mockChatResponse,mockEmbedResponse,mockListModelsResponse). - Conversation Store: Transactional STM-backed
InMemoryStoreandConversationStoretypeclass. - New API Endpoints:
Ollama.API.Embed(/api/embed),Ollama.API.Blobs(/api/blobs),Ollama.API.Ps(/api/ps),Ollama.API.Version(/api/version). - Benchmark Suite: Criterion/tasty-bench suite in
bench/Main.hsmeasuring serialization and throughput.
Changed
- MonadIO / MonadUnliftIO Polymorphism: All API functions use
MonadIO m =>/MonadUnliftIO m =>signatures instead of dual*Mvariants. - Typed Newtypes:
ModelName,Digest,Base64Image,Duration,Versionreplace primitive string types. - Unified Error Type:
OllamaErrorsum type with structured constructors andExceptioninstance.
Deprecated
embeddingsendpoint (/api/embeddings) marked deprecated in favor of/api/embed.