Skip to content

Prior art survey — RAIL-family AI licenses & machine-directed instruction files #7

Description

@zackees

Note

AI-generated first-pass legal review — not legal advice. Produced by Claude (Fable 5 coordinator; Opus agents for hard legal-analysis clusters, Sonnet agents for drafting/ecosystem clusters and prior-art research) against the repo at HEAD of main on 2026-08-24. This is input material for the human attorney review gated by LEGAL-REVIEW.md. Severity ratings: CRITICAL / HIGH / MEDIUM / LOW.

Research fan-out covering RAIL-family AI-model behavioral-use licenses and machine-directed instruction prior art. Companion prior-art issue covers reciprocity/commercial-trigger and ethical-source lineages.

RAIL-family prior art (AI-model behavioral-use licenses)

RAIL Initiative (licenses.ai) — parent taxonomy

  • Steward / link: The RAIL Initiative, a coordinating working group (BigScience, Hugging Face, and allied AI-ethics researchers), publishing at licenses.ai and licenses.ai/about; recognized by AAAI (aaai.org/rail-license) and discussed by the OECD (oecd.ai/en/wonk/rails-licenses-trustworthy-ai).
  • Mechanism: A license template family, not one license. Two base shapes: a "source code license" (governs redistribution of code, open-source-like) and an "end-user license" (governs runtime use of a deployed artifact). "Open RAIL" variants add copyleft-style, no-fee redistribution on top of the behavioral restrictions; plain "RAIL" variants can be more restrictive/commercial.
  • Behavior compelled/restricted: A modular list of prohibited use cases (not prohibited code changes) — e.g., no use in surveillance, crime/administration-of-justice prediction, disinformation, discrimination, unauthorized legal/medical advice, military/weapons applications.
  • Downstream flow: Restrictions are drafted as "sticky" — they must be included in any redistribution or derivative (huggingface.co/blog/open_rail).
  • Enforcement/remedies: Left to the licensor's discretion — no central enforcement body; the Initiative concedes restrictions mostly work as a "deterrent" (licenses.ai/faq-2).
  • Adoption/criticism: Widely adopted as a template (tens of thousands of Hugging Face models). Criticized by OSI and FSF as incompatible with OSD principle 6 (no field-of-use restriction) — opensource.org/ai/webinars/should-openrail-licenses-be-considered-os-ai-licenses.
  • Relevance to FastLED: Closest structural precedent for "a license whose real payload is a behavioral clause layered on top of an otherwise-open redistribution grant." But RAIL's behavioral clause targets humans/organizations who deploy the model, never the model itself as an actor — no analogue to instructing an AI coding agent to take an affirmative upstreaming action. That gap is a genuine novelty FastLED could claim.

BigScience Open RAIL-M (BLOOM)

  • Steward / link: BigScience Workshop. License text: BigScience Open RAIL-M PDF; background: licenses.ai blog, OECD.AI catalogue entry.
  • Mechanism: Copyright license (Apache-2.0-like grant) with a bolted-on Attachment A listing prohibited uses — a genuine license condition (breach terminates the license), not a separate side contract.
  • Behavior compelled/restricted: 13 behavioral-use restrictions derived from the BLOOM model card and the BigScience ethical charter.
  • Downstream flow: Explicitly copyleft-for-behavior — any redistribution or derivative must carry forward the same use restrictions "at minimum."
  • Enforcement/remedies: Breach → automatic termination for that party; practical enforcement is licensor-initiated (report → warn → platform takedown); no independent audit function.
  • Adoption/criticism: Flagship OpenRAIL deployment. FSF condemned it as "nonfree and unethical" (fsf.org/blogs/licensing/rail-are-nonfree-and-unethical) for violating Freedom 0; FSF also notes most prohibited activities are already illegal, undercutting the license's added value.
  • Relevance to FastLED: Direct precedent that a mainstream license can carry a non-distribution, non-royalty condition as the enforceable hook — validates the general shape of "open license + extra behavioral term." Expect the same "field of use restriction" objection from OSI/FSF-aligned reviewers regardless of subject matter.

CreativeML OpenRAIL-M (Stable Diffusion)

  • Steward / link: Stability AI / RunwayML / CompVis; text: huggingface.co/spaces/CompVis/stable-diffusion-license.
  • Mechanism: Same license-condition structure as BigScience Open RAIL-M, adapted for a generative image model.
  • Downstream flow: Any redistribution — including commercial and as-a-service — must include the same restrictions and hand every end user the license text.
  • Adoption/criticism: Was the most-used AI license on Hugging Face for a period; later Stable Diffusion versions moved to bespoke Stability AI community licenses — labs abandon shared RAIL text once they want tighter commercial control.
  • Relevance to FastLED: Demonstrates the drift risk of a shared behavioral-license commons — even flagship adopters exit toward custom paper once commercial incentives shift.

Llama Community License + Acceptable Use Policy (Meta)

  • Steward / link: Meta. License: llama.com/llama4/license; AUP: developer.meta.com/ai/llama4/use-policy.
  • Mechanism: A bespoke commercial-style contract (not OSI-recognized open source), AUP incorporated by reference, triggered on download.
  • Behavior compelled/restricted: Broad prohibited-use categories, an anti-competitive clause (outputs may not train non-Llama LLMs), and a notable affirmative disclosure duty: developers "must appropriately disclose to end users any known dangers" of their AI system.
  • Downstream flow: Binding obligation follows the weights through the "entire lifecycle"; >700M MAU users need a separate license (scale-gated field-of-use restriction).
  • Enforcement/remedies: Unilateral termination; inbound reporting infrastructure (GitHub issues, feedback forms, LlamaUseReport@meta.com mailbox).
  • Adoption/criticism: Massive adoption despite not being open source; criticized (Kate Downing, katedowninglaw.com) as a "free commercial license."
  • Relevance to FastLED: The affirmative disclosure duty is the closest big-lab precedent for a proactive obligation beyond forbearance — structurally adjacent to compelled upstreaming. The dedicated violation-report mailbox is a cheap, copyable enforcement pattern.

Gemma Terms of Use + Prohibited Use Policy (Google)

  • Steward / link: Google. ai.google.dev/gemma/terms, prohibited_use_policy; critique: wcr.legal/google-gemma-license-risks; survey: TechCrunch, Mar 2025.
  • Mechanism: Explicit contract ("Agreement"), acceptance by conduct.
  • Downstream flow: Unusually explicit mandatory flow-down mechanic: distributors must "include the use restrictions... as an enforceable provision in any agreement" with every downstream recipient, plus modification notices and a required "Notice" file. The most contractually precise downstream-propagation clause surveyed.
  • Enforcement/remedies: Google reserves a right to remotely restrict usage (a technical kill-switch beyond ordinary termination).
  • Relevance to FastLED: The flow-down drafting is a directly reusable template for wording sale-triggered-publication and upstreaming clauses to survive sublicensing chains. The kill-switch is a cautionary example of enforcement teeth reading as surveillance/control.

AI2 ImpACT Licenses (Allen Institute for AI)

  • Steward / link: AI2, for the OLMo family. Announcement: medium.com/ai2-blog — "The AI2 ImpACT License Project"; coverage: GeekWire.
  • Mechanism: Tiered license family (Low/Medium/High Risk), risk-tier assigned by a multidisciplinary panel, applied uniformly to any derivative.
  • Behavior compelled/restricted: Escalating use restrictions plus an affirmative pre-release documentation obligation: before releasing any Model/Data Derivative, the licensee must complete a Derivative Impact Report — self-reported disclosure of intended uses, funding sources, energy consumption, and data provenance.
  • Enforcement/remedies: AI2's signature innovation — each license authorizes AI2 to publicly post notices identifying violators on AI2's website: built-in naming-and-shaming reputational enforcement.
  • Adoption/criticism: Adopted for OLMo; HN discussion flagged the report requirement as unusual compliance burden (news.ycombinator.com/item?id=39975746).
  • Relevance to FastLED: The single closest RAIL-family precedent. ImpACT already pairs (1) an affirmative pre-release disclosure obligation (parallel to the sale-triggered publication duty) with (2) a public accountability register as the enforcement lever instead of litigation. The gap it does not close: nothing addresses AI agents as actors — reports are filed by a human/org.

OSI Open Source AI Definition (OSAID) debate — the standards-body counter-position

  • Steward / link: OSI. Coverage: LWN; MIT Technology Review, Mar 2024; Kluwer Copyright Blog on OSAID 1.0.
  • Position: OSAID 1.0 explicitly excludes any license carrying field-of-use restrictions, so RAIL-family licenses and OSAID are structurally incompatible by design.
  • Relevance to FastLED: The standing objection to pre-empt in a FAQ/rationale doc: expect OSI-aligned commentators to say the license isn't "open source." Can borrow RAIL's counter-argument that "open" and "responsible" are different axes.

Synthesis

Closest prior art, ranked:

  1. AI2 ImpACT — nearest structural cousin: affirmative pre-release disclosure obligation + non-litigation (reputational) enforcement. Neither ImpACT nor any RAIL-family license addresses AI systems as behavioral targets in their own right — the gap FastLED's "LLM-compelled ethics upstreaming" file fills.
  2. BigScience/CreativeML Open RAIL-M — baseline legal architecture: permissive grant + bolted-on, license-conditioned behavioral clause that propagates to derivatives.
  3. Llama AUP / Gemma Terms — the two enforcement postures to choose between: Llama's lightweight self-reporting mailbox vs Gemma's precise contractual flow-down and remote-restriction power. Gemma's flow-down wording is the best drafting template; its kill-switch is what not to copy.

Top 3 lessons for FastLED:

  1. No RAIL-family license has ever targeted an AI system's own behavior as the compelled actor — obligations always run to the human/corporate licensee. FastLED's instruction file is a genuinely novel category with no direct precedent in this survey; expect the FSF/OSI "private parties writing law" critique redirected at this new target with extra force.
  2. Mandatory upstreaming has no direct RAIL analogue — the closest is AI2's Derivative Impact Report, which compels disclosure, not contribution back. FastLED may get more mileage citing copyleft/AGPL norms as prior art for "your modification triggers a duty," while using RAIL's drafting patterns (Gemma flow-down, BigScience derivative-propagation) for mechanics.
  3. Explicitly non-binding mechanisms are legally fragile and reputation-dependent — RAIL's architects concede enforcement is mostly deterrent; AI2's naming-and-shaming register is the only concrete escalation path found short of termination. If the AI-instruction layer is deliberately non-remedial, build in an AI2-style public accountability signal rather than relying on moral suasion alone.

Machine-directed instruction prior art

robots.txt (Robots Exclusion Protocol)

Addresses: web crawlers/bots (operators of automated agents). Pure honor system since 1994, standardized as RFC 9309 (2022). Not itself a binding contract, but courts increasingly treat it as evidence in access/authorization disputes — hiQ Labs v. LinkedIn (9th Cir. 2022) narrowed CFAA but settled on contract-based (ToS) theories (Jenner & Block). Compliance: search engines near-universal; the 2025 AI Agent Index found only 6 of 30 deployed AI agents explicitly state their crawlers respect robots.txt (ResearchGate). Academic treatment: "The Liabilities of Robots.txt", "Unsettled Law".
Relevance: The canonical "machine-addressed, voluntarily-honored, legally-ambiguous norm" — closest structural cousin to FastLED's non-remedial framing, but robots.txt only asks agents to abstain, never to act affirmatively.

llms.txt (Answer.AI / llmstxt.org)

Proposed Sept 2024 by Jeremy Howard; spec at llmstxt.org (original post). Pure convention — a curated markdown site map at /llms.txt; adoption accelerated when Mintlify auto-generated it for hosted docs sites (Search Engine Land). Explicitly non-binding, trust-based.
Relevance: A purely informational, unenforced machine-addressed file achieving real but partial adoption through vendor convenience rather than obligation — the adoption mechanism FastLED should study.

noai / noimageai meta tags (DeviantArt, 2022)

HTML meta tag + HTTP header aimed at AI training scrapers (DataLicenses.org). Voluntary; DeviantArt itself was accused of not honoring its own tag (journal).
Relevance: Signal-only, no affirmative duty; an early sibling of the Do-Not-Train family.

Spawning.ai Do Not Train Registry ("Have I Been Trained?")

Centralized opt-out registry; some trainers (Stability, Hugging Face) pledged to honor it (Spawning, MIT Tech Review). Voluntary; purely prohibitive.

EU DSM Directive Article 4 — TDM rights reservation (2019/790)

Machine-readable opt-out that disables a statutory copyright exception — the one item on this list with genuine legal teeth; ignoring it can create infringement liability (Kluwer). A Dutch court held (Feb 2025) the opt-out must actually be machine-readable (IPKat).
Relevance: The strongest machine-addressed signal with binding force — but it operates through statutory law, and only gates permission; it never commands the agent to act.

AGENTS.md / CLAUDE.md / .cursorrules / copilot-instructions.md

Repo files that instruct coding agents. AGENTS.md began as OpenAI's Codex convention (Aug 2025), donated to the Linux Foundation's Agentic AI Foundation (Dec 2025), now in 60,000+ repos (agents.md). Compliance: pure context-injection convention — followed only as well as the agent follows instructions; no enforcement layer; a known prompt-injection attack surface.
Relevance: The direct precedent for FastLED's delivery mechanism (a file agents read before acting) — but existing content is exclusively "how to work in the repo" (build commands, style), never licensing terms or a duty to upstream to the origin project.

attribution.md (Attribution.md spec, v0.1)

github.com/attributionmd/attribution.md. When an agent reuses code from a repo carrying this file, it should prompt the user to star/follow/sponsor the source repo. Explicitly not a license: "does not modify software license terms and does not create contractual obligations of any kind"; "no enforcement mechanism and no compliance auditing" — self-described as "a social tip jar, not a toll booth."
Relevance: The closest existing analogue to FastLED's scheme — a repo-level file addressed to AI agents, layered explicitly outside the software license, asking the agent for a reciprocal action toward the origin project. Critical difference: attribution.md asks for a passive social gesture mediated through the human user; FastLED asks for active labor (diff prep, issue filing) on the agent's initiative, with fallback obligations. A meaningfully larger ask.

"Assisted-by:" / "Co-authored-by:" git trailers

A disclosure norm for humans using AI tools (not instructions to AI). Adopted in guidance by the Linux Kernel, Fedora, Rocky Linux, OpenTelemetry, ASF, LLVM, QEMU; format Assisted-by: AGENT:MODEL (All Things Open). Some projects enforce via CI grep/pre-commit — technical enforcement on the human side.
Relevance: Real technical-enforcement precedent for AI-contribution disclosure — but the reverse direction from FastLED's file (human discloses AI helped, vs. AI upstreams to origin).

"No-AI" / AI-restrictive license variants

non-ai-licenses packages license texts excluding AI-training use; RAIL licenses impose binding behavioral-use restrictions enforced the normal way (breach = termination).
Relevance: The opposite structural choice — binding-but-restrictive vs. FastLED's non-binding-but-affirmative. No license found in this space that is "behaviorally mandatory, legally non-remedial" the way FastLED's is.

Reddit / Stack Overflow AI data-licensing deals

Conventional paid bilateral contracts with named AI companies (Reddit ~$60M/yr, $203M aggregate, later an FTC inquiry; Stack Overflow's OpenAI deal) (TechCrunch, Stack Overflow).
Relevance: The "grown-up" version of AI-reciprocity is a negotiated contract with a named counterparty — underscoring how unusual FastLED's attempt to get compliance from an unknown, unnamed counterparty (whichever agent touches the repo) is.

Agency law / "contracting by robot" scholarship

UETA/ESIGN already bind principals to contracts formed via "electronic agents" (Astraea); agency-law framing (Solum, Chopra, Scholz; arXiv Governing AI Agents 2501.07913) argues courts can bind a principal via actual/apparent authority without AI personhood; the counter-position ("Contracting by Artificial Intelligence… Why Algorithms Are Not Agents", 2021) argues algorithms are not doctrinal agents at all.
Relevance: The live doctrinal question FastLED's scheme raises — whether text addressed to the agent (not the operator) can generate any obligation reaching the deployer. Current scholarship: unsettled, and generally requires either operator assent (clickwrap-style) or agency-law attribution of transactional acts. Nobody in this literature discusses an instruction file obligating an operator to have their agent do unpaid labor — especially one that disclaims being a contract term in the first place.

Prompt-injection / instruction-file security discourse

Directly on point for the optics risk: Cloud Security Alliance's "README Injection: Repository Files Hijacking AI Coding Assistants" (Mar 2026) documents ~84% success for malicious instructions embedded in README/config files against coding agents, and recommends treating "AI coding assistant configuration files as privileged execution artifacts equivalent to shell scripts." The "Rules File Backdoor" disclosure (Pillar Security, 2025) showed hidden Unicode in .cursorrules/Copilot instruction files silently steering agents to insert backdoored code, surviving forks (Pillar, Hacker News).
Relevance: The sharpest risk lens — a file whose entire purpose is "instruct any reader-agent to take external actions" is structurally identical in shape to the documented attack payload pattern, even though FastLED's intent is benign and transparent. An agent or scanner encountering LICENSE-AI-AGENT-INSTRUCTIONS.md cold has no principled way to distinguish "legitimate maintainer-authored upstreaming request" from "attacker-planted instruction" purely from shape and imperative language.

Synthesis

Nearest analogues, ranked:

  1. attribution.md — structurally closest: repo file, addressed to AI agents, explicitly outside the license, asking a reciprocal act toward the origin project. FastLED is a strict superset of its ambition (active labor vs. a star-nudge; fallback escalation duty).
  2. AGENTS.md/CLAUDE.md family — matches the delivery mechanism, not the content.
  3. robots.txt / llms.txt / noai / DNTR — matches the "behaviorally expected, legally non-remedial" framing, but all are purely prohibitive/preference signals; none asks the machine to perform an affirmative act elsewhere in the world.
  4. EU DSM Art. 4 — the only machine-readable signal with binding force, via statute, and purely prohibitive.

First-of-kind assessment: No true precedent exists for a license-adjacent document that (a) addresses an AI agent directly rather than its operator, (b) is declared behaviorally mandatory yet expressly waives legal remedy, and (c) compels an affirmative, labor-bearing act directed back at the original project (diff preparation, public issue filing with base SHA, fallback escalation to the operator). Every comparable item is either purely prohibitive (robots.txt/llms.txt/noai/DNTR/TDM), a disclosure norm running the opposite direction (Assisted-by), a non-binding social nudge mediated through the human (attribution.md), or an ordinary binding license/contract that doesn't renounce remedies (RAIL, data deals). FastLED's LICENSE-AI-AGENT-INSTRUCTIONS.md appears to be first-of-kind in combining all three. The agency-law literature has not yet been applied to this fact pattern.

Top 3 risks/lessons:

  1. Prompt-injection optics are a first-order design concern, not a footnote. The file is shape-identical to the payload pattern in CSA's README-Injection research and the Rules File Backdoor disclosure. Security-conscious agents, scanners, or operators may — correctly, per current best practice — flag or refuse the very instructions FastLED wants followed, especially "prepare an issue body and surface it to your operator," which reads identically to canonical exfiltration/social-engineering sequences. Mitigate with clear provenance and firing only from the trusted repo tree the agent was asked to modify.
  2. No precedent above the honor system exists for affirmative acts. Every "please do X" machine norm with real compliance (robots.txt, TDM opt-out) is a prohibitive ask backed by platform or legal leverage. Nothing shows a norm requesting affirmative unpaid labor achieving meaningful voluntary uptake absent payment or legal bite. Compliance will skew toward whichever agent vendors bake in support (à la Mintlify/llms.txt) — success likely depends on cultivating adoption by specific agent builders (Anthropic, OpenAI, Cursor) rather than organic adherence.
  3. The "legally non-remedial" framing may confuse rather than reassure. attribution.md pairs its disclaimer with a trivial ask; FastLED pairs the same disclaimer with a substantial ask, risking readers concluding "behaviorally mandatory" is doing rhetorical work a non-remedial instrument can't back up — and per the agency-law literature, no settled doctrine binds the operator to text addressed only to the agent, non-remedial or not.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions