Repository navigation
Add artifacts - #404
Add artifacts#404simonpcouch wants to merge 23 commits into
Conversation
|
Preview deployed to Connect ( Deployed from commit b3e2414. |
|
Preview root: https://posit-dev.github.io/commons/pr-404/ Python site preview: https://posit-dev.github.io/commons/pr-404/py/ Built from the latest commit on this branch. The R links in it point at the published R site, which no pull request rebuilds. |
|
Preview deployed to Connect ( Deployed from commit b3e2414. |
| text should stand on its own. | ||
|
|
||
| {% if has_edit_artifact %} | ||
| ## Documents |
There was a problem hiding this comment.
This really should be a skill. We should refactor. The trust system tool into a skill and then support a few built-in skills, starting with that one and this "artifact" one.
There was a problem hiding this comment.
The tricky part is that this would also require (?) deferred tool loading... for. theedit_artifact() tool.
Here's one description og how that could possibly work:
Reading through this, I'm wondering whether the execute-style approach is how we ought to go. i.e.
- MCP tools and any built-in tools we flag (currently, the notebook tools) are not part of the normal tool list, and live behind a meta-tool execute
- The names and description snippets for each of the tools gated behind the meta-tool are provided in the system prompt
- When the model is interested in calling one of them, it calls loadTool to load the full description and parameter description of the tool, then invokes it via execute
- The "normal" tool UI for those tools can be "lifted" into the tool UI for execute
There are many judgments to be made here, though, and I feel on-the-fence for many of them.
|
Review from Claude Opus 5.5 Agent-written review (assistant-review skill)Review: PR #404 — Add artifacts | branch
|
|
Review from GPT-6.1 Sol Agent-written reviewReviewed
|
|
|
||
| Start the document on its own line with `<artifact id="sales-by-region" title="Sales by region">` and end it with `</artifact>`. Between them goes the full `.qmd` source: YAML frontmatter, then Markdown and R code cells. The id is a short lowercase slug. Writing the tag again with the same id replaces that document; a new id makes a new one. In the chat, the document is replaced by a link to it, so don't repeat its contents in your reply. | ||
|
|
||
| In the frontmatter, set `title` and, if useful, `subtitle` or `date`; the format, theme, and execution options are set for you. |
There was a problem hiding this comment.
This is minor, but I kept running into situations were the model does not write any YAML frontmatter, and so the artifact render fails.
Maybe it would help to strengthen the wording here or include an example of what an example <artifact> should look like.
|
Looks great, this is basically how I envisioned artifacts working! I think the trusted code results -> CSV approach for reproducibility is probably good at this stage (maybe in the future commons layers will be more portable). Some thoughts:
Some issues I ran into:
|


Closes #71.
My goal here was:
reports.mov
It also felt important to me that the artifact streamed in. (Notably, this is not how artifacts work on claude.ai anymore, surprisingly. Will have to think about why.) In addition, I wanted computations to happen eagerly—e.g. when the agent completes an inline
ror a code cell, it begins evaluating then rather than when the whole report is done streaming in. This means that we're not actually using Quarto here, instead using a hand-rolled knitr worker. (More context on that at the top ofartifact-knit.R.)The agent can edit the report after, a la Canvas.
Tried to make the abstractions usable on the Python side, but does not implement in Python.
Deferring Quarto dashboards / Shiny apps, deferring "Share to Connect".