From 75ecd679ba1753c268ada8c94383daeba57e1c91 Mon Sep 17 00:00:00 2001 From: Jung Do Hyun Date: Tue, 25 Aug 2026 12:07:24 +0900 Subject: [PATCH] =?UTF-8?q?docs:=20'=EB=88=84=EA=B0=80=20=EC=8D=A8?= =?UTF-8?q?=EC=95=BC=20=ED=95=98=EB=82=98'=20=EC=A0=88=20=EC=B6=94?= =?UTF-8?q?=EA=B0=80=20=E2=80=94=20=EC=A0=81=ED=95=A9=C2=B7=EB=B6=80?= =?UTF-8?q?=EC=A0=81=ED=95=A9=20=EB=8C=80=EC=83=81=20(3=EA=B0=9C=20?= =?UTF-8?q?=EC=96=B8=EC=96=B4)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- README.ja.md | 31 +++++++++++++++++++++++++++++++ README.ko.md | 28 ++++++++++++++++++++++++++++ README.md | 32 ++++++++++++++++++++++++++++++++ 3 files changed, 91 insertions(+) diff --git a/README.ja.md b/README.ja.md index 32a8808..ee1fc27 100644 --- a/README.ja.md +++ b/README.ja.md @@ -38,6 +38,37 @@ OpenAI・Anthropic互換APIを提供し、プロンプト・レスポンス・ Anthropicの`/v1/messages`ネイティブ対応 + プレフィックスキャッシング(TTFT実測0.82秒 → 0.14秒)により、Claude Codeがそのまま接続できる。 +## 誰が使うべきか + +一文で言えば: **コードを外部に送れない、トークンを大量に消費する、あるいはGPUを集中的に +使う**開発者。 + +**適しているケース:** + +1. **プライバシー制約のある開発者** — 会社の規定や規制で外部LLM APIにコードを送れない + 場合。推論・プロンプト・使用量統計まで自分のAWSアカウント内で完結し、全体がOSSなので + 監査可能。このグループにとって代替手段は商用APIではなく「何もない」である。 +2. **利用上限に縛られたヘビーパワーユーザー** — エージェントのファンアウトで1日100M+ + トークンを消費する場合。時間定額(月約$100-200)はトークン無制限であり、この規模から + 同一モデルAPIと同等、商用ファーストパーティAPI比では数倍安くなる。128Kコンテキストの + 定額利用もティア課金APIに対する構造的な利点。 +3. **バッチワークロード運用者** — 夜間の大量コード分析・生成のようにGPUを埋めて使う + 場合。稼働率が上がるとトークン単価が3-5倍改善する、スポットの最適ポイント。 +4. **AWSクレジット・コミット(EDP)保有組織** — API費用は別枠の支出だが、GPUスポットは + 既存のAWSコミットメント内で消化できる。 + +**適していないケース:** + +- **ライトユーザー** — 月数十Mトークン以下なら、同じオープンウェイトモデルを提供する + サードパーティAPIの方が圧倒的に安い。この領域でtoken-forgeを選ぶ理由はプライバシー + だけである。 +- **無停止が必要なサービス** — スポット回収時に数分の再確保ギャップを許容する設計(R3) + のため、SLAのあるプロダクションサービングには不向き。 +- **最上位モデルの品質が必須なワークロード** — オープンウェイト30B-355B級で不十分なら + 適さない。 +- AWSアカウント・クォータ管理自体が負担な場合 — 初回利用にスポットクォータの引き上げ + 申請が必要。 + ## アーキテクチャ ```mermaid diff --git a/README.ko.md b/README.ko.md index f09c515..b223ef6 100644 --- a/README.ko.md +++ b/README.ko.md @@ -36,6 +36,34 @@ 잡힌 곳만 남긴다(First-Acquired-Wins). Anthropic `/v1/messages` 네이티브 + prefix caching(TTFT 실측 0.82초 → 0.14초)으로 Claude Code가 그대로 붙는다. +## 누가 써야 하나 + +한 문장으로: **코드를 외부로 못 보내거나, 토큰을 아주 많이 쓰거나, GPU를 몰아서 쓰는** +개발자. + +**적합한 경우:** + +1. **프라이버시 제약이 있는 개발자** — 회사 규정·규제로 외부 LLM API에 코드를 못 보내는 + 경우. 추론·프롬프트·사용량 통계까지 자기 AWS 계정 안에서 완결되고 전체가 OSS라 감사 + 가능하다. 이 그룹에게 대안은 상용 API가 아니라 "아무것도 없음"이다. +2. **한도에 묶인 헤비 파워유저** — 에이전트 팬아웃으로 일 100M+ 토큰을 태우는 경우. + 시간 정액(월 약 $100-200)은 토큰 무제한이라, 이 물량부터 동일 모델 API와 동률이 되고 + 상용 1차 API 대비 수 배 저렴해진다. 128K 컨텍스트 정액도 티어 과금 API 대비 이점. +3. **배치 워크로드 운영자** — 야간 대량 코드 분석·생성처럼 GPU를 채워 쓰는 경우. + 가동률이 오르면 토큰 단가가 3-5배 개선되는 스팟의 최적 지점이다. +4. **AWS 크레딧·커밋(EDP) 보유 조직** — API 비용은 별도 지출이지만 GPU 스팟은 기존 + AWS 커밋 안에서 소진된다. + +**맞지 않는 경우:** + +- **라이트 유저** — 월 수십M 토큰 이하면 동일 오픈 웨이트 모델을 서빙하는 서드파티 + API가 압도적으로 싸다. 이 구간에서 token-forge를 고를 이유는 프라이버시뿐이다. +- **무중단이 필요한 서비스** — 스팟 회수 시 수 분의 재확보 공백을 감수하는 설계(R3)라 + SLA가 있는 프로덕션 서빙에는 부적합하다. +- **최상위 모델 품질이 필수인 워크로드** — 오픈 웨이트 30B-355B급으로 부족하다면 맞지 + 않다. +- AWS 계정·쿼터 관리 자체가 부담인 경우 — 첫 진입에 스팟 쿼터 증설 신청이 필요하다. + ## 아키텍처 ```mermaid diff --git a/README.md b/README.md index e40628d..e131f44 100644 --- a/README.md +++ b/README.md @@ -43,6 +43,38 @@ evidence and numbers. Anthropic `/v1/messages` support plus prefix caching (measured TTFT dropping from 0.82s to 0.14s) means Claude Code connects without any adjustment. +## Who is this for + +In one sentence: developers who **cannot send code outside, consume very large token +volumes, or run GPU workloads in concentrated bursts**. + +**A good fit if you are:** + +1. **A developer under privacy constraints** — company policy or regulation forbids + sending code to external LLM APIs. Inference, prompts, and even usage statistics stay + inside your own AWS account, and the whole stack is open source and auditable. For + this group the alternative is not a commercial API but "nothing at all". +2. **A heavy power user hitting subscription limits** — agent fan-out burning 100M+ + tokens a day. Flat GPU-hours (about $100-200/month) mean unlimited tokens: from this + volume on, self-hosting matches same-model API pricing and beats first-party commercial + APIs by multiples. Flat-rate 128K context also avoids context-tiered API billing. +3. **A batch workload operator** — nightly bulk code analysis or generation that fills + the GPU. Higher utilization improves token unit cost by 3-5x, the sweet spot for spot. +4. **An organization with AWS credits or committed spend (EDP)** — API bills are separate + spend, but GPU spot burns down your existing AWS commitment. + +**Not a good fit if you:** + +- **Use it lightly** — below tens of millions of tokens a month, third-party APIs serving + the same open-weight models are far cheaper. In that range the only reason to choose + token-forge is privacy. +- **Need uninterrupted service** — the design accepts a few minutes of re-acquisition gap + on spot reclaim (R3); it is not for SLA-bound production serving. +- **Require frontier-model quality** — if open-weight 30B-355B models are not enough, + this is not the tool. +- Find AWS account and quota management itself a burden — the first run requires a spot + quota increase request. + ## Architecture ```mermaid