Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions README.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,37 @@ OpenAI・Anthropic互換APIを提供し、プロンプト・レスポンス・
Anthropicの`/v1/messages`ネイティブ対応 + プレフィックスキャッシング(TTFT実測0.82秒 →
0.14秒)により、Claude Codeがそのまま接続できる。

## 誰が使うべきか

一文で言えば: **コードを外部に送れない、トークンを大量に消費する、あるいはGPUを集中的に
使う**開発者。

**適しているケース:**

1. **プライバシー制約のある開発者** — 会社の規定や規制で外部LLM APIにコードを送れない
場合。推論・プロンプト・使用量統計まで自分のAWSアカウント内で完結し、全体がOSSなので
監査可能。このグループにとって代替手段は商用APIではなく「何もない」である。
2. **利用上限に縛られたヘビーパワーユーザー** — エージェントのファンアウトで1日100M+
トークンを消費する場合。時間定額(月約$100-200)はトークン無制限であり、この規模から
同一モデルAPIと同等、商用ファーストパーティAPI比では数倍安くなる。128Kコンテキストの
定額利用もティア課金APIに対する構造的な利点。
3. **バッチワークロード運用者** — 夜間の大量コード分析・生成のようにGPUを埋めて使う
場合。稼働率が上がるとトークン単価が3-5倍改善する、スポットの最適ポイント。
4. **AWSクレジット・コミット(EDP)保有組織** — API費用は別枠の支出だが、GPUスポットは
既存のAWSコミットメント内で消化できる。

**適していないケース:**

- **ライトユーザー** — 月数十Mトークン以下なら、同じオープンウェイトモデルを提供する
サードパーティAPIの方が圧倒的に安い。この領域でtoken-forgeを選ぶ理由はプライバシー
だけである。
- **無停止が必要なサービス** — スポット回収時に数分の再確保ギャップを許容する設計(R3)
のため、SLAのあるプロダクションサービングには不向き。
- **最上位モデルの品質が必須なワークロード** — オープンウェイト30B-355B級で不十分なら
適さない。
- AWSアカウント・クォータ管理自体が負担な場合 — 初回利用にスポットクォータの引き上げ
申請が必要。

## アーキテクチャ

```mermaid
Expand Down
28 changes: 28 additions & 0 deletions README.ko.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,34 @@
잡힌 곳만 남긴다(First-Acquired-Wins). Anthropic `/v1/messages` 네이티브 + prefix
caching(TTFT 실측 0.82초 → 0.14초)으로 Claude Code가 그대로 붙는다.

## 누가 써야 하나

한 문장으로: **코드를 외부로 못 보내거나, 토큰을 아주 많이 쓰거나, GPU를 몰아서 쓰는**
개발자.

**적합한 경우:**

1. **프라이버시 제약이 있는 개발자** — 회사 규정·규제로 외부 LLM API에 코드를 못 보내는
경우. 추론·프롬프트·사용량 통계까지 자기 AWS 계정 안에서 완결되고 전체가 OSS라 감사
가능하다. 이 그룹에게 대안은 상용 API가 아니라 "아무것도 없음"이다.
2. **한도에 묶인 헤비 파워유저** — 에이전트 팬아웃으로 일 100M+ 토큰을 태우는 경우.
시간 정액(월 약 $100-200)은 토큰 무제한이라, 이 물량부터 동일 모델 API와 동률이 되고
상용 1차 API 대비 수 배 저렴해진다. 128K 컨텍스트 정액도 티어 과금 API 대비 이점.
3. **배치 워크로드 운영자** — 야간 대량 코드 분석·생성처럼 GPU를 채워 쓰는 경우.
가동률이 오르면 토큰 단가가 3-5배 개선되는 스팟의 최적 지점이다.
4. **AWS 크레딧·커밋(EDP) 보유 조직** — API 비용은 별도 지출이지만 GPU 스팟은 기존
AWS 커밋 안에서 소진된다.

**맞지 않는 경우:**

- **라이트 유저** — 월 수십M 토큰 이하면 동일 오픈 웨이트 모델을 서빙하는 서드파티
API가 압도적으로 싸다. 이 구간에서 token-forge를 고를 이유는 프라이버시뿐이다.
- **무중단이 필요한 서비스** — 스팟 회수 시 수 분의 재확보 공백을 감수하는 설계(R3)라
SLA가 있는 프로덕션 서빙에는 부적합하다.
- **최상위 모델 품질이 필수인 워크로드** — 오픈 웨이트 30B-355B급으로 부족하다면 맞지
않다.
- AWS 계정·쿼터 관리 자체가 부담인 경우 — 첫 진입에 스팟 쿼터 증설 신청이 필요하다.

## 아키텍처

```mermaid
Expand Down
32 changes: 32 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,38 @@ evidence and numbers.
Anthropic `/v1/messages` support plus prefix caching (measured TTFT dropping from 0.82s to
0.14s) means Claude Code connects without any adjustment.

## Who is this for

In one sentence: developers who **cannot send code outside, consume very large token
volumes, or run GPU workloads in concentrated bursts**.

**A good fit if you are:**

1. **A developer under privacy constraints** — company policy or regulation forbids
sending code to external LLM APIs. Inference, prompts, and even usage statistics stay
inside your own AWS account, and the whole stack is open source and auditable. For
this group the alternative is not a commercial API but "nothing at all".
2. **A heavy power user hitting subscription limits** — agent fan-out burning 100M+
tokens a day. Flat GPU-hours (about $100-200/month) mean unlimited tokens: from this
volume on, self-hosting matches same-model API pricing and beats first-party commercial
APIs by multiples. Flat-rate 128K context also avoids context-tiered API billing.
3. **A batch workload operator** — nightly bulk code analysis or generation that fills
the GPU. Higher utilization improves token unit cost by 3-5x, the sweet spot for spot.
4. **An organization with AWS credits or committed spend (EDP)** — API bills are separate
spend, but GPU spot burns down your existing AWS commitment.

**Not a good fit if you:**

- **Use it lightly** — below tens of millions of tokens a month, third-party APIs serving
the same open-weight models are far cheaper. In that range the only reason to choose
token-forge is privacy.
- **Need uninterrupted service** — the design accepts a few minutes of re-acquisition gap
on spot reclaim (R3); it is not for SLA-bound production serving.
- **Require frontier-model quality** — if open-weight 30B-355B models are not enough,
this is not the tool.
- Find AWS account and quota management itself a burden — the first run requires a spot
quota increase request.

## Architecture

```mermaid
Expand Down
Loading