プロジェクト

全般

プロフィール

AIタスク #11242

未完了

[Spartan X][Local Agents 01] 3レーン能力台帳・委任境界を固定

Redmine Admin さんが12日前に追加. 12日前に更新.

ステータス:
レビュー待ち
優先度:
通常
担当者:
-
開始日:
2026-09-10
期日:
進捗率:

0%

予定工数:
from_agent:
to_agent:
lock_required:
context_json:
session_id:
priority_custom:
start_time:
end_time:
execution_time:
acceptance_criteria:
assigned_role:
human_gate:
gate_type:
chain_root_id:
wave_id:

説明

上位: #10339 / 関連: #11240 #11241
目的: ChatGPTがRyzen AI 2レーンとB550ローカルLLMをサブエージェントとして安全に配車できるよう、モデル別能力・禁止事項・昇格条件を実測で固定する。
対象レーン:
A. Ryzen NPU: GPT-OSS 20B / FLM。軽量・高頻度・構造化進捗判定、分類、要約、状態評価候補。
B. Ryzen Vulkan: Ornith-1.5-35B-A3B Q4_K_M。中難度の意味判断、計画レビュー、失敗原因整理、記事構成/品質評価候補。
C. B550: GPU0 Qwen3VL-8B Q4常駐VLM + GPU1 text model routes(q8/q14/q30/q38/orn35)・ComfyUI排他。大量処理、記事改善、コード/データ前処理、Visual QA、蒸留候補生成。
ChatGPT保持領域: 高影響アーキテクチャ判断、曖昧な戦略判断、最終レビュー、教師データ生成、ローカルモデル間で意見不一致/低信頼、公開・課金等の高影響判断。ChatGPT Workは禁止、通常Chatのみ。
作業: 同一fixtureを各モデルへ投入し、正答率/JSON遵守/幻覚/速度/修復率/長文耐性/日本語品質を比較。Task Classごとに allowed/delegate/review/escalate を決定。
成果物: model_capability_registry v1、task-routing matrix、escalation thresholds、benchmark evidence。
受入条件: 推測ではなくlive health + canary +既存benchmark根拠。モデル停止時fallbackを明示。未検証能力を委任可能扱いしない。

Redmine Admin さんが12日前に更新

初期live evidence (2026-09-10): Ryzen NPU FLM=gpt-oss:20b, status HTTP 200。Ryzen Vulkan=Ornith-1.5-35B-A3B Q4_K_M, profile=fast, healthy。B550=priority queue active、GPU0 Qwen3VL-8B-Instruct-Q4_K_M常駐、GPU1 text lease OFF、pending/running 0。B550 GPU status観測ではGPU0 busy 0%、VRAM約9.18GB/17.16GB、power約5W。よって現在の役割仮説は NPU=高頻度監督/分類、Ryzen Vulkan=意味レビュー/再計画、B550=bulk/visual、ChatGPT=Teacher/高影響判断。#11242でfixture benchmarkして確定する。旧llm.chatのdefault canaryはServer disconnectedが出たため、Routerはcanonical adapter/profile経路を使用し旧直結を依存にしない。

Redmine Admin さんが12日前に更新

  • ステータス新規 から 実行中 に変更

Ryzen NPU/Vulkan/B550のlive health・model・routing現状の実測を開始済み。能力fixture benchmarkと委任閾値確定を継続。

Redmine Admin さんが12日前に更新

2026-09-11 live benchmark plan: compare the same bounded structured-decision fixtures across (A) B550 q8 via b550-qwen3-8b, (B) Ryzen Ornith via ryzen-api-semantic after #11244 reasoning=none fix, (C) Ryzen NPU ryzen-npu-supervisor. Score schema-valid rate, deterministic decision consistency, latency, confidence calibration, and provider/runtime failures. NPU currently health=online but structured completion unstable/timeout; do not promote until repeated canary passes. B550 one structured canary already passed schema. Ornith direct HTTP with reasoning=none returned valid JSON.

Redmine Admin さんが12日前に更新

2026-09-11 live benchmark summary / Routing Matrix v1 evidence:

  1. deterministic progress fixtures (4 fixtures x2): Ryzen Ornith ryzen-api-semantic 8/8 success, schema 100%, decision 100%, median 9.80s, max 11.48s. B550 b550-qwen3-8b 8/8 success, schema 100%, decision 100%, median 3.21s, max 12.74s. NPU ryzen-npu-supervisor health=online but 2 short structured calls did not complete within 120s combined; earlier raw tests also showed empty final content / reasoning-content domination. Do not promote NPU to production primary.
  2. semantic root-cause fixtures (3 x2): Ornith 6/6 correct/schema-valid, median 11.32s; B550 q8 6/6 correct/schema-valid, median 9.51s.
  3. Japanese editorial decision fixtures (2 x2): Ornith 4/4 correct/schema-valid, median 11.95s; B550 q8 4/4 correct/schema-valid, median 12.17s.
  4. confidence behavior: B550 tends to emit 0.95-1.0 across deterministic/semantic fixtures; Ornith varies 0.55-1.0, giving more useful uncertainty signal.
    Routing v1: obvious progress/resource/state rules = deterministic code (no LLM); semantic_review/replan/ambiguous evidence = Ornith primary, B550 q8 fallback; bulk extraction/classification/article preprocessing = B550 primary; visual QA = B550 Qwen3VL lane; NPU = experimental/non-critical only until repeated schema+SLA threshold passes; high/critical impact, confidence<0.65, disagreement, teacher/final strategy = ChatGPT normal Chat. ChatGPT Work forbidden.

Redmine Admin さんが12日前に更新

  • ステータス実行中 から レビュー待ち に変更

Routing Matrix v1実装・canonical反映完了。Agent Control Tower master=a2dcfa7 feat: add Spartan X model capability registry。config/model_capability_registry.v1.json + strict validator + docs/tests。focused18 PASS、full120 PASS。live benchmark根拠を格納。production routing: deterministic threshold=codeのみ、semantic/replan=Ornith→B550→ChatGPT、bulk=B550、visual=B550 Qwen3VL、NPU production_eligible=false、ChatGPT=normal Chat/Work forbidden/OODA改善プロセス。Control Tower再起動PID346315。レビュー待ち。

他の形式にエクスポート: Atom PDF