Skip to content

Latest commit

 

History

82 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llm

A multi-provider LLM client for Carp. Supports Anthropic, OpenAI, Ollama, and Google Gemini behind a single common API.

Built on http-client and json.

Installation

(load "git@github.com:carpentry-org/llm@0.5.1")

Requires OpenSSL for HTTPS providers (Anthropic, OpenAI, Gemini). Ollama over plain HTTP works without OpenSSL.

Usage

Basic chat (Ollama, no API key)

(let [config (LLM.ollama "http://localhost:11434")
      req (LLM.chat-request "llama3" [(Message.user "hello")] 256 0.7)]
  (match (LLM.chat &config &req)
    (Result.Success r) (println* (LLMResponse.content &r))
    (Result.Error e) (IO.errorln &(LLMError.str &e))))

Switching providers

The same LLM.chat works against any provider — just change the config constructor:

(LLM.anthropic "sk-ant-...")     ; Anthropic
(LLM.openai    "sk-...")          ; OpenAI
(LLM.ollama    "http://...")      ; Ollama
(LLM.gemini    "AIza...")         ; Gemini

Streaming

(match (LLM.chat-stream &config &req)
  (Result.Success stream)
    (do
      (while-do true
        (match (LlmStream.poll &stream)
          (Maybe.Nothing) (break)
          (Maybe.Just tok) (IO.print &tok)))
      (LlmStream.close stream))
  (Result.Error e) (IO.errorln &(LLMError.str &e)))

Streaming with tool calls

Use poll-event instead of poll to receive both text and tool calls from a streaming response. Incremental tool call data (OpenAI/Anthropic) is accumulated internally and emitted as complete ToolCall values:

(let [tools [(ToolDef.init @"get_weather" @"Get weather" schema)]
      req (LLM.chat-request-with-tools "gpt-4"
            [(Message.user "Weather in Paris?")] 256 0.7 tools)]
  (match (LLM.chat-stream &config &req)
    (Result.Success stream)
      (do
        (while-do true
          (match (LlmStream.poll-event &stream)
            (Maybe.Nothing) (break)
            (Maybe.Just evt)
              (match-ref &evt
                (StreamEvent.Text tok) (IO.print tok)
                (StreamEvent.ToolCallEvent tc)
                  (println* "tool: " (ToolCall.name tc)
                            " args: " (ToolCall.arguments tc)))))
        (LlmStream.close stream))
    (Result.Error e) (IO.errorln &(LLMError.str &e))))

Tool use

(let [schema (JSON.obj [(JSON.entry @"type" (JSON.Str @"object"))
                        (JSON.entry @"properties"
                          (JSON.obj [(JSON.entry @"city"
                            (JSON.obj [(JSON.entry @"type" (JSON.Str @"string"))]))]))])
      tools [(ToolDef.init @"get_weather" @"Get current weather" schema)]
      req (LLM.chat-request-with-tools "gpt-4"
            [(Message.user "Weather in Paris?")] 256 0.7 tools)]
  (match (LLM.chat &config &req)
    (Result.Success r)
      (when (> (Array.length (LLMResponse.tool-calls &r)) 0)
        (let [tc (Array.unsafe-nth (LLMResponse.tool-calls &r) 0)]
          (println* "calling " (ToolCall.name tc) " with " (ToolCall.arguments tc))))
    (Result.Error e) (IO.errorln &(LLMError.str &e))))

To send a tool result back, build a follow-up request including the assistant's tool call message and a tool result:

(let [msgs [(Message.user "Weather in Paris?")
            (Message.from-response &r)
            (Message.tool-result &tool-call-id "22°C, sunny")]
      req2 (LLM.chat-request-with-tools "gpt-4" msgs 256 0.7 @&tools)]
  (LLM.chat &config &req2))

Agentic tool-use loop

LLM.chat-loop handles the full call-handle-respond cycle automatically:

(let [config (LLM.ollama "http://localhost:11434")
      tools [(ToolDef.init @"get_weather" @"Get weather" schema)]
      msgs [(Message.user "Weather in Paris?")]
      handler (fn [tc]
                (if (= (ToolCall.name tc) "get_weather")
                  @"22C, sunny"
                  @"unknown tool"))]
  (match (LLM.chat-loop &config "llama3" msgs 256 0.7 &tools handler 10)
    (Result.Success r) (println* (LLMResponse.content &r))
    (Result.Error e) (IO.errorln &(LLMError.str &e))))

The handler receives each ToolCall by reference and returns the result as a String. The loop repeats until the model stops calling tools or the iteration limit is reached.

JSON output

; Plain JSON mode (any valid JSON)
(LLM.chat-request-json model msgs max-tokens temp)

; Schema-constrained JSON
(LLM.chat-request-with-schema model msgs max-tokens temp schema)

Anthropic has no native JSON mode — for them, the library falls back to a system prompt instruction. Best-effort, not guaranteed. All other providers use their native JSON mode.

Embeddings

(let [config (LLM.openai "sk-...")
      req (LLM.embedding-request "text-embedding-3-small" [@"hello" @"world"])]
  (match (LLM.embed &config &req)
    (Result.Success r)
      (println* "got " (Array.length (EmbeddingResponse.embeddings &r)) " embeddings")
    (Result.Error e) (IO.errorln &(LLMError.str &e))))

Anthropic does not offer an embeddings API — calling LLM.embed with an Anthropic config returns a Transport error.

Retries

LLM.chat, LLM.chat-stream, LLM.chat-loop and LLM.embed make exactly one attempt. Each has a -with-retry counterpart that takes a RetryPolicy and retries 429 rate limits, the 500, 502, 503 and 504 server errors, Anthropic's 529, and transport failures:

(let [config (LLM.openai "sk-...")
      req (LLM.chat-request "gpt-4" [(Message.user "hello")] 256 0.7)]
  (match (LLM.chat-with-retry &config &req &(RetryPolicy.default))
    (Result.Success r) (println* (LLMResponse.content &r))
    (Result.Error e) (IO.errorln &(LLMError.str &e))))

RetryPolicy.default is three attempts with a 500ms base delay doubling to a 30s cap. Every other 4xx comes straight back to you, since retrying a bad request only wastes quota. retryable-statuses names the set, so a provider with its own transient status is a field away.

A transport failure means a connection that was refused, reset or closed. A failure that another attempt cannot change — a malformed base URL, a URL without a host, a redirect chain that could not be followed, a certificate that does not verify, or a response that arrived but could not be parsed — comes straight back to you on the first attempt. Requests carry no read timeout, so an endpoint that accepts the connection and then stalls blocks the attempt instead of failing it, and no retry follows.

A retry re-sends the whole request, so a generation that the provider already ran and billed can run again: the client cannot tell a connection that failed on the way out from one that failed after the provider had answered.

A Retry-After response header, in either the delta-seconds or the HTTP-date form, replaces the computed delay — still clamped to max-delay-ms, so a server cannot park you indefinitely. Set honour-retry-after to false to ignore it.

Delays are deterministic by default. Set jitter to true to draw each computed delay uniformly from [0, delay] instead, which keeps many clients sharing one provider from retrying in lockstep.

; five attempts, 200ms base, 10s cap, Retry-After honoured, jittered
(RetryPolicy.init 5 200 10000 true true [429 500 502 503 504 529])

; the default policy, also retrying 409 conflicts
(RetryPolicy.set-retryable-statuses (RetryPolicy.default)
                                    [409 429 500 502 503 504 529])

; never retry, which is what the plain entry points use
(RetryPolicy.none)

For streaming, the policy covers the initial response only: once a stream is handed back, a failure mid-stream is yours to handle. For chat-loop it applies per request, not to the loop as a whole.

API

Construction

Function Purpose
LLM.anthropic key ProviderConfig for Anthropic
LLM.openai key ProviderConfig for OpenAI
LLM.ollama base-url ProviderConfig for Ollama
LLM.gemini key ProviderConfig for Gemini
LLM.chat-request model msgs max-tokens temp Plain text request
LLM.chat-request-with-tools model msgs max-tokens temp tools Request with tool definitions
LLM.chat-request-json model msgs max-tokens temp JSON output mode
LLM.chat-request-with-schema model msgs max-tokens temp schema Schema-constrained JSON
LLM.embedding-request model input Embedding request (input is an array of strings)
Message.user content User message
Message.assistant content Assistant message
Message.system content System message
Message.tool-result call-id content Tool result message (for follow-ups)
Message.from-response &r Convert an LLMResponse to an assistant message for history

Sending

Function Purpose
LLM.chat config req Synchronous chat. Returns (Result LLMResponse LLMError)
LLM.chat-loop config model msgs max-tokens temp tools handler max-iters Agentic tool-use loop. Calls chat, invokes handler per tool call, repeats until done or limit reached
LLM.chat-stream config req Streaming chat. Returns (Result LlmStream LLMError)
LlmStream.poll stream Returns (Maybe String) — next text token, or Nothing when done
LlmStream.poll-event stream Returns (Maybe StreamEvent) — text or tool call event, or Nothing when done
LLM.embed config req Generate embeddings. Returns (Result EmbeddingResponse LLMError)
LLM.chat-with-retry config req policy chat, retrying per the RetryPolicy
LLM.chat-loop-with-retry config model msgs max-tokens temp tools handler max-iters policy chat-loop, retrying each request
LLM.chat-stream-with-retry config req policy chat-stream, retrying the initial response
LLM.embed-with-retry config req policy embed, retrying per the RetryPolicy
LlmStream.close stream Close the underlying connection

Errors

(deftype LLMError
  (Transport [String])          ; connection / DNS / network errors
  (Api [Int String String]))    ; HTTP status, error type, message

LLMError.str &e formats either variant for display. API errors are parsed from each provider's specific error JSON shape.

Provider quirks

  • Anthropic: System messages are extracted from the message array and sent in a separate system field. No native JSON mode. No embeddings API.
  • OpenAI: Standard format. System messages are kept as a "system" role in the messages array.
  • Ollama: Same shape as OpenAI for messages and tool calls. NDJSON for streaming (one JSON object per line) instead of SSE.
  • Gemini: Different shape entirely (contents/parts instead of messages). Uses the /v1beta endpoint for tool support. Streaming uses :streamGenerateContent?alt=sse.

The library hides all of this — you just call LLM.chat and the provider's build/parse functions translate.

Testing

carp -x test/llm.carp

The unit tests don't make network calls. For live tests against real APIs, see examples/anthropic.carp, examples/openai.carp, examples/ollama.carp, examples/gemini.carp and their _stream, _tool, _json variants, plus examples/embeddings.carp for the embeddings API. Most require an API key in the appropriate environment variable.


Have fun!

About

a llm integration for carp (includes openai, anthropic, gemini, and ollama)

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors