dsprrr lets you write LLM features in R as small programs instead of prompt strings. You declare a task’s inputs and typed outputs; dsprrr builds the prompt, calls the model through ellmer, and returns an R list with the types you asked for. Once you have labeled examples, you can score the program with a metric and let an optimizer tune its instructions and few-shot examples against that score. The design follows DSPy.
If you have a prompt that already works and no data to measure it against, plain ellmer is enough. dsprrr earns its keep when you want typed outputs across many inputs, a score you can track, or a prompt tuned on examples instead of by hand.
dsprrr is not on CRAN yet. Install the development version from GitHub:
# install.packages("pak")
pak::pak("JamesHWade/dsprrr")You also need credentials for a model provider, for example
OPENAI_API_KEY in your .Renviron.
A signature names the inputs and outputs of a task:
library(dsprrr)
signature(
"review -> sentiment: enum('positive', 'negative', 'neutral'), stars: int, summary: string"
)
#>
#> ── Signature ──
#>
#> ── Inputs
#> • review: "string" - Input: review
#>
#> ── Output
#> Type: "object(sentiment: enum(positive, negative, neutral), stars: integer,
#> summary: string)"
#>
#> ── Instructions
#> Given the fields `review`, produce the fields `sentiment`, `stars`, `summary`.module() turns it into something you can run with any ellmer chat:
chat <- ellmer::chat_openai(model = "gpt-6-luna")
analyzer <- module(signature(
"review -> sentiment: enum('positive', 'negative', 'neutral'), stars: int, summary: string"
))
result <- run(
analyzer,
review = "I've been using this blender for 6 months now. It's incredibly powerful and easy to clean. The only downside is it's quite loud. Overall, I'm very happy with it.",
.llm = chat
)
str(result)
#> List of 3
#> $ sentiment: chr "positive"
#> $ stars : int 4
#> $ summary : chr "Powerful and easy-to-clean blender, but a bit loud."That output was recorded from a real call to gpt-4.1 in the
structured outputs
tutorial;
gpt-6-luna may word the summary differently.
To measure and improve a module, give it labeled rows and a metric:
scores <- evaluate(
analyzer,
labeled_reviews,
metric = metric_exact_match(field = "sentiment"),
.llm = chat
)
scores$mean_score
optimized <- analyzer |>
compile(
BootstrapFewShot(metric = metric_exact_match(field = "sentiment")),
trainset = labeled_reviews,
.llm = chat
)| Area | Functions |
|---|---|
| Define tasks | signature(), input(), with_instructions() |
| Run them | module(), run(), run_dataset(), run_async(), run_stream() |
| Other ways to answer | chain_of_thought(), react(), best_of_n(), refine(), ensemble(), program_of_thought(), code_act() |
| Compose | pipeline(), %>>%, module_fn() |
| Measure | evaluate(), metric_exact_match(), metric_f1(), vitals bridges |
| Optimize | compile() with LabeledFewShot(), BootstrapFewShot(), MIPROv2(), GEPA(), COPRO(), SIMBA() and more; optimize_grid() |
| Inspect | get_last_prompt(), inspect_history(), summarize_traces(), session_cost() |
| Save | save_program(), pin_module_config() |
Experimental: rlm_module() (a model explores a large R object by
writing code), flex() (GEPA rewrites a whole program),
with_decisions() and ReAnchor() (calibrated decisions), Omni() and
agentic optimization harnesses.
The documentation site has six tutorials that start from a first call (Tutorial 1), how-to guides, concept articles, and a guide for DSPy users.
Experimental. The API may change. See the open issues for the roadmap.
