1. はじめに
アクションルーターインテリジェンス理論に関する関連論文は、ルーティングはアクション制御の問題であり、テキスト分類の問題ではないという理論的基礎を確立しました。正式な定義 R: (コンテキスト × 意図 × 状態) → アクションは、キーワードとセマンティック ルーティングを包含しながら、構成的なゲートの統合と責任の保持を可能にすることが証明されました。しかし、実装のない理論は建物のない建築です。
This paper bridges the gap. We present the complete Action Router implementation as deployed within MARIA OS, covering three questions that the theory paper left open. First, how is the routing triple (Context, Intent, State) constructed from raw system inputs? The Intent Parser layer answers this by defining concrete extraction pipelines for each component. Second, how does the Action Resolver select among potentially thousands of registered actions in real time? We answer with a hierarchical search algorithm that exploits MARIA coordinate structure for O(log|A|) action selection. Third, how does the system improve over time? We formalize recursive self-improvement as an online learning problem and prove convergence guarantees.
1.1 Architecture Overview
The Action Router consists of three layers, each with a clearly defined interface:
| Layer | Name | Input | Output | Latency Budget |
|---|---|---|---|---|
| L1 | Intent Parser | Raw input x + session metadata | (Context C, Intent I) | ≤ 12ms |
| L2 | Action Resolver | (C, I, S) triple | Action a ∈ A_feasible | ≤ 10ms |
| L3 | Gate Controller | Action a + risk assessment | Gated execution envelope | ≤ 5ms |
合計レイテンシ バジェットは 27 ミリ秒で、P99 の目標である 30 ミリ秒に対して 3 ミリ秒の余裕が残っています。各層は独立して展開可能であり、水平方向に拡張可能です。これらのレイヤーは、MARIA OS SDK で定義された型付きインターフェイスを介して通信し、あるレイヤーへの変更が他のレイヤーを破壊しないようにします。
1.2 設計原則
3 つの原則が実装の指針となります。(1) 分類なし、制御のみ — どのレイヤーもカテゴリ ラベルを生成しません。すべての出力は構造化データまたは実行可能なアクションのいずれかです。 (2) デフォルトでフェールクローズ — いずれかのレイヤーが信頼性の高い出力を生成できない場合、リクエストはデフォルトのハンドラーにルーティングされるのではなく、人間によるレビューにエスカレーションされます。 (3) 構築によって観察可能 — すべてのレイヤーが、入出力ペア、信頼スコア、タイミング データを含む構造化テレメトリを発行し、再帰学習ループを可能にします。
2. Layer 1: Intent Parser
2.1 Context Extraction Pipeline
The Intent Parser constructs the Context C from three sources: (a) the user’s MARIA coordinate and authority level, retrieved from the authentication session; (b) the session history, including prior routing decisions and their outcomes within the current interaction; (c) active organizational policies that may constrain available actions (e.g., a freeze on refunds during audit periods). Context extraction is deterministic and rule-based, requiring no ML inference:
Context extraction completes in under 2ms because all inputs are available from in-memory caches (session cache, policy cache, MARIA coordinate registry).
2.2 Intent Extraction Model
Intent extraction transforms raw text x and context C into a structured intent I. Unlike keyword extraction, intent extraction produces a four-field structure:
The goal field g ∈ G is drawn from a finite goal taxonomy specific to the MARIA Universe. Goals are not keywords — they are structured specifications. For example, goal = resolve_multi_issue(issues=[contract_cancellation, refund_processing, compliance_verification]) specifies a compound goal with three sub-components. The constraint set restricts which actions are acceptable (e.g., “refund must not exceed $10,000”). Priority is a continuous score reflecting the user’s stated or inferred priority. Urgency is a discrete field determined by context signals (SLA timers, escalation history, user role).
2.3 軽量の意図分類子
インテント抽出モデルは意図的に軽量化されており、組織データに基づいて微調整された 1,200 万のパラメーターを備えた 2 層トランスフォーマー エンコーダーです。 (a) レイテンシー: 7B パラメーター モデルは推論に 40 ~ 80 ミリ秒を必要とし、レイテンシー バジェット全体を消費するため、意図の抽出には大規模な言語モデルを避けます。 (b) 決定論: 大規模なモデルは、監査ログを複雑にする非決定的な出力を示します。 (c) 組織の特異性: 50,000 のラベル付き組織リクエストに基づいて微調整された小規模モデルは、ドメイン固有のインテント抽出において汎用の大規模モデルよりも優れたパフォーマンスを発揮します。
インテント分類器は、単一の GPU で 8ms の P99 推論レイテンシで、保持された組織データに対して 94.2% の精度を達成します。 2 ミリ秒のコンテキスト抽出と組み合わせると、レイヤー 1 は余裕を持って 12 ミリ秒の予算内で完了します。
2.4 曖昧さの検出と明確化
When the intent classifier’s confidence falls below a threshold τ = 0.7, the Intent Parser triggers a clarification protocol rather than guessing. The clarification protocol generates a structured question based on the top-2 candidate intents:
このフェールクローズ動作により、不確実なルーターが間違った推測をしたときに発生するエラーの連鎖が防止されます。運用環境では、明確化率はリクエストの 6.8% であり、明確化されたリクエストは 99.1% のルーティング精度を達成しています。
3. Layer 2: Action Resolver
3.1 Action Registry
The Action Resolver operates on a pre-registered action space A. Each action is registered with its full specification: preconditions, effects, responsibility assignment, gate level, and cost function. The registry is organized hierarchically by MARIA coordinate:
ActionRegistry
G1 (Galaxy: Enterprise)
U1 (Universe: Sales)
P1 (Planet: Customer Success)
Z1 (Zone: Retention)
a_101: initiate_retention_offer
a_102: process_cancellation
a_103: escalate_to_manager
Z2 (Zone: Billing)
a_201: process_refund
a_202: adjust_invoice
U2 (Universe: Legal)
P1 (Planet: Contracts)
Z1 (Zone: Review)
a_301: initiate_contract_review
a_302: flag_compliance_issueThis hierarchical structure enables O(log|A|) action lookup by narrowing the search path: Galaxy → Universe → Planet → Zone → Actions. For a typical enterprise with |A| = 500, the search visits at most 4 levels × 5 nodes per level = 20 nodes, compared to 500 for a flat linear scan.
3.2 Precondition Filtering
Given the routing triple (C, I, S), the Action Resolver first computes the feasible action set by evaluating preconditions:
ここで、A_scope は、ユーザーの権限レベルによって決定される MARIA 座標スコープ内のアクションのサブセットです。座標 G1.U1.P1.Z1 を持つユーザーは、そのゾーン (昇格されたアクセス許可が付与されている場合は親スコープ) に登録されているアクションにのみアクセスできます。前提条件の評価は、スレッド プールを使用して実行可能な候補セット全体で並列化され、20 ~ 50 アクションの一般的なスコープ サイズの場合、3 ミリ秒未満で完了します。
3.3 効果に基づくアクションのランキング
リゾルバーは、実行可能なアクションの中で、予測される効果の質、つまりアクションの予測結果がユーザーが指定した目標にどの程度一致するかによって候補をランク付けします。
距離関数 d は、テキストの埋め込みではなく、構造化された状態表現に作用します。たとえば、目標が 3 つのサブ問題を含む replace_multi_issue である場合、 d は、アクションの予測効果によって対処されるサブ問題の数を、緊急度によって重み付けしてカウントします。この構造化された距離の計算により、意味的には似ているが操作的には異なるアクションが同様のスコアを受け取る埋め込み混同問題が回避されます。
3.4 Compound Action Composition
When no single action satisfies a compound intent, the resolver composes multiple actions into an action plan:
Composition uses a greedy set-cover algorithm: at each step, select the action whose effect covers the most unsatisfied goal components. The greedy algorithm achieves a (1 - 1/e) approximation ratio for submodular goal coverage, which is provably optimal in polynomial time. In practice, compound intents require 2-3 actions on average, and the composition completes in under 2ms.
4. Layer 3: Gate Controller
4.1 Risk Assessment
The Gate Controller receives the selected action (or action plan) and computes a risk score that determines the gate level. The risk assessment combines three factors:
ImpactScore measures the magnitude of the action’s effects (financial amount, number of affected entities, scope of state change). ReversibilityScore measures how easily the action can be undone (fully reversible = 0, partially reversible = 0.5, irreversible = 1.0). ConfidenceGap measures the uncertainty in the routing decision (difference between the Action Resolver’s confidence in the top-ranked action versus the second-ranked action). A high ConfidenceGap indicates the router is uncertain, warranting additional oversight.
4.2 ゲートレベルの割り当て
リスク スコアは、構成可能なしきい値を通じてゲート レベルにマップされます。
| Risk Score | Gate Level | Execution Mode | Expected Latency |
|---|---|---|---|
| [0, 0.3) | Level 0: Auto-Execute | Immediate execution, async audit log | 0ms (fire-and-forget) |
| [0.3, 0.6) | Level 1: Soft Review | Execute immediately, flag for retrospective review | 0ms + async review |
| [0.6, 0.8) | Level 2: Human Review | Queue for human approval before execution | Minutes to hours |
| [0.8, 1.0] | Level 3: Escalation | Route to senior decision-maker with full context bundle | Hours to days |
しきい値は MARIA Universe ごとに構成できるため、さまざまなビジネス ユニットがリスク許容度を調整できます。高リスクのトレーディング デスクはレベル 2 のしきい値を 0.8 (積極的) に設定し、コンプライアンス部門はレベル 2 のしきい値を 0.4 (保守的) に設定する場合があります。
4.3 実行エンベロープの構築
ゲート コントローラーは、アクションを実行エンベロープ (実行と監査に必要なものすべてを含む構造化パケット) にラップします。
The execution envelope is immutable once constructed. It is persisted to the MARIA OS audit log before the action is dispatched. The TTL (time-to-live) field ensures that gated actions that are not approved within a configurable window are automatically expired and the requester is notified, preventing stale routing decisions from executing in a changed system state.
5. 再帰的な自己改善
5.1 The Feedback Loop
The Action Router improves continuously through a feedback loop that connects execution outcomes back to routing weights. After every action completes (or fails), the outcome is recorded:
The outcome record captures whether the action succeeded, how well it satisfied the original intent, how long it took, and whether it produced unexpected side effects. This outcome data feeds into two learning mechanisms: weight updates for the Action Resolver and threshold calibration for the Gate Controller.
5.2 Online Weight Updates
The Action Resolver maintains a weight vector w ∈ ℝ^{|A|} that biases action selection. After observing outcome o_t, the weights are updated using the exponentiated gradient algorithm:
ここで、ℓ_t(a) は時間 t でのアクション a によって発生する損失 (ゴール距離、コスト、失敗ペナルティを組み合わせたもの)、η は学習率、Z_t は正規化定数です。この乗法更新には、次の 3 つの望ましい特性があります。(a) どのアクションにも重みを 0 に割り当てることはありません (探索は保存されます)。 (b) 後から考えると、レート O(√T) で最良の固定アクションに収束します。 (c) 計算的には自明です (更新ごとに O(|A|))。
5.3 収束解析
定理 1 (再帰的改善収束)。 学習率 η = √(ln|A| / (2T)) による累乗勾配更新では、アクション ルーターの累積損失は次の条件を満たします。
|A| の場合= 500、T = 30,000 (1 日あたり 1,000 件の決定で約 30 日の実稼働ルーティング)、決定ごとの平均後悔は √(2 · 30000 · ln 500) / 30000 ≈ 0.018 です。これは、ルーターが 30 日後に最適な固定ポリシーの 1.8% 以内にあることを意味します。これは、観測された精度が 93.4% から 97.8% に +4.4% 向上したことと一致しています。
5.4 Gate Threshold Calibration
ゲート コントローラーのリスクしきい値も再帰学習によって調整されます。私たちはベイジアン アプローチを使用します。各しきい値の事前値は組織のポリシーによって設定され、事後値は観察された偽陽性 (不必要にゲートされたアクション) と偽陰性 (ゲートされていないアクションが害を引き起こす) 率に基づいて更新されます。
非対称重み α_FP および α_FN は、偽陽性 (不必要な遅延) と偽陰性 (制御されていないリスク) の相対コストを反映します。実際には、α_FN ≫ α_FP は 5 ~ 10 倍です。これは、システムが保守的であることを意味します。つまり、真のリスクを見逃さないように、不必要なゲートを許容します。
6. Scaling Architecture for 100+ Agent Deployments
6.1 The Scaling Challenge
Enterprise MARIA OS deployments can involve 100+ concurrent agents across multiple Universes, each with its own action space. Naive centralized routing creates a bottleneck: all routing decisions funnel through a single resolver that must search the entire action space. At 10,000 requests per second and |A| = 2,000, the centralized approach exceeds the 30ms latency target.
6.2 Coordinate-Based Sharding
Action Router を MARIA 座標でシャーディングします。各ユニバースは、その組織スコープ内のすべてのルーティング決定を処理する専用のルーティング パーティションを受け取ります。
Each partition maintains its own action registry, weight vector, and gate thresholds. Cross-Universe routing (rare, approximately 3% of requests) is handled by a lightweight meta-router that determines the target Universe before delegating to the appropriate partition.
6.3 階層型アクションキャッシュ
各パーティション内で、3 レベルのアクション キャッシュを維持します。
- L1 Cache (Zone-local): The 10 most frequently selected actions per Zone, stored in-memory with sub-microsecond access. Cache hit rate: 72%.
- L2 キャッシュ (プラネットローカル): プラネットごとに最も頻繁に選択された 50 個のアクション。共有メモリ領域に保存されます。キャッシュヒット率:91%(L1との累計)。
- L3 Cache (Universe-wide): The full action registry for the Universe, stored in a Redis cluster. Cache hit rate: 100% (by definition).
キャッシュ階層により、一般的なケースでは平均アクション ルックアップが 5 ミリ秒 (完全なレジストリ検索) から 0.8 ミリ秒 (L1 ヒット) に短縮されます。キャッシュの無効化はイベント駆動型です。アクションの前提条件が変更されると (エージェントがオフラインになるなど)、関連するキャッシュ エントリが MARIA OS イベント バスを介してただちに無効化されます。
6.4 Throughput Analysis
With 4 Universe partitions, each handling 2,500 rps, and the cache hierarchy reducing per-request latency to 14ms (P50), the system sustains 10,000 rps with P99 latency of 28ms. The bottleneck shifts from action search to intent extraction (Layer 1), which we address by deploying multiple intent classifier replicas behind a load balancer.
7. Integration with the Decision Pipeline State Machine
7.1 The Decision Pipeline
The MARIA OS Decision Pipeline implements a 6-stage state machine:
proposed → validated → [approval_required | approved] → executed → [completed | failed]MARIA OS におけるすべての決定は、このパイプラインを通過します。ルーティングされたアクションによって決定が作成されるため、アクション ルーターはパイプラインとインターフェイスする必要があります。ルーターによって選択されたアクションは「提案された」状態でパイプラインに入り、実行前に適切なステージを通過する必要があります。
7.2 プロダクトオートマトン
We formalize the integration as a product automaton of the Action Router state and the Decision Pipeline state. Let Q_R = {idle, parsing, resolving, gating, dispatched} be the Action Router states and Q_P = {proposed, validated, approval_required, approved, executed, completed, failed} be the Pipeline states. The product automaton Q = Q_R × Q_P has |Q_R| × |Q_P| = 5 × 7 = 35 states, of which 18 are reachable:
網羅的な列挙によって、(MARIA OS valid_transitions テーブルで定義されている) 12 個の有効なパイプライン遷移すべてがルーティング層から到達可能であることを検証します。これは、アクション ルーターがパイプラインを通じて有効な決定を実行できることを意味します。デッド ステートは存在しません。
7.3 遷移マッピング
各ルーティング結果は、特定のパイプライン遷移にマップされます。
| Router Outcome | Pipeline Transition | Condition |
|---|---|---|
| Action selected, gate = L0 | proposed → validated → approved → executed | Auto-execute path |
| Action selected, gate = L1 | proposed → validated → approved → executed | Execute + async review |
| Action selected, gate = L2 | proposed → validated → approval_required | Queue for human approval |
| Action selected, gate = L3 | proposed → validated → approval_required | Escalate to senior |
| No feasible action | proposed → failed | Fail-closed |
| Ambiguous intent | (no pipeline entry) | Clarification loop |
7.4 アトミック性とロールバック
製品オートマトンは、ルーティングとパイプラインの遷移がアトミックであることを保証します。ルーティングからパイプラインへのハンドオフ全体が成功するか、システムがルーティング前の状態にロールバックします。これは、2 フェーズ コミット プロトコルを使用して実装されます。フェーズ 1 では、実行エンベロープを構築し、パイプライン スロットを予約します。フェーズ 2 では、アクションをディスパッチし、パイプラインの状態を進めます。フェーズ 2 が失敗した場合 (ターゲット エージェントが利用できないなど)、フェーズ 1 がロールバックされ、更新された状態でルーティングの決定が再試行されます。
8. Production Metrics and Benchmarks
8.1 導入構成
The benchmarks are conducted on a simulated production deployment with the following configuration: 4 MARIA Universes (Sales, Legal, Compliance, Operations), 12 Planets, 48 Zones, 127 active agents, 523 registered actions, and an average request rate of 10,000 routing decisions per second during peak hours. The evaluation period is 30 days with a total of 8.6 million routing decisions.
8.2 Accuracy Over Time
| Day | Routing Accuracy | Clarification Rate | Fail-Closed Rate |
|---|---|---|---|
| 1 | 93.4% | 6.8% | 1.2% |
| 7 | 95.1% | 5.4% | 0.9% |
| 14 | 96.3% | 4.7% | 0.7% |
| 21 | 97.2% | 4.1% | 0.5% |
| 30 | 97.8% | 3.6% | 0.4% |
The accuracy improvement follows the theoretical O(√T) convergence rate. The clarification rate decreases as the intent classifier improves through fine-tuning on production data. The fail-closed rate decreases as the action registry expands to cover edge cases identified during operation.
8.3 Latency Profile
| Component | P50 | P90 | P99 | P99.9 |
|---|---|---|---|---|
| Intent Parser (L1) | 7ms | 10ms | 12ms | 18ms |
| Action Resolver (L2) | 4ms | 7ms | 10ms | 15ms |
| Gate Controller (L3) | 2ms | 3ms | 5ms | 8ms |
| Total (end-to-end) | 14ms | 22ms | 28ms | 38ms |
The P99 total of 28ms is within the 30ms target. The P99.9 of 38ms exceeds the target but occurs at a rate of 1 in 1,000 requests, acceptable for enterprise workloads where critical requests receive priority queue treatment.
8.4 Scaling Efficiency
| Metric | 1 Partition | 2 Partitions | 4 Partitions | 8 Partitions |
|---|---|---|---|---|
| Max RPS | 3,200 | 6,100 | 10,400 | 18,700 |
| P99 Latency | 28ms | 27ms | 28ms | 29ms |
| Scaling Efficiency | 1.0x | 0.95x | 0.81x | 0.73x |
8 つのパーティションでは、パーティション間のルーティング オーバーヘッド (ユニバースの境界を越えるリクエストの 3%) により、スケーリング効率が低下します。ほとんどのエンタープライズ展開では、4 つのパーティションで、ほぼ線形のスケーリングで十分なスループットが提供されます。
8.5 ベースラインとの比較
| Metric | Keyword Router | Semantic Router | Action Router (Day 1) | Action Router (Day 30) |
|---|---|---|---|---|
| Accuracy | 62.1% | 74.3% | 93.4% | 97.8% |
| P99 Latency | 12ms | 89ms | 28ms | 26ms |
| Gate Compliance | 71.3% | 76.8% | 99.5% | 99.7% |
| Audit Completeness | 41.2% | 55.1% | 100% | 100% |
| Responsibility Attribution | 34.8% | 48.3% | 97.1% | 98.4% |
9. ディスカッション
9.1 実装から得た教訓
Three implementation lessons stand out. First, the intent classifier’s training data quality matters more than model size. A 12M parameter model trained on 50,000 high-quality organizational examples outperforms a 7B parameter general model on domain-specific intents by 11 percentage points. This is because organizational intents have distributional properties (e.g., compound goals, authority-dependent semantics) that general models have not been trained on.
次に、アクション レジストリは生きた成果物であり、継続的なキュレーションが必要です。 30 日間の評価を通じて、23 の新しいアクションが登録され、8 つのアクションが廃止され、41 のアクションの前提条件が更新されました。再帰学習ループは、既存のアクションと高い信頼度で一致しないリクエストのクラスターを検出することで、新しいアクションの必要性を特定します。
第三に、フェールクローズ設計はユーザーの信頼を構築するために不可欠です。導入の初期段階では、6.8% の解明率が弱点 (「システムが私を理解していない」) として認識されていました。自動ルーティングされたリクエストの 93.4% に対して、明確化されたリクエストでは 99.1% のルーティング精度が達成されたことをユーザーが観察した後、明確化のプロンプトは肯定的なシグナル (「システムが慎重である」) になりました。 30 日目までに、明確化されたリクエストのユーザー満足度スコアは、自動ルーティングされたリクエストのユーザー満足度スコアを 8 ポイント上回りました。
9.2 Limitations
Action Router には 3 つの既知の制限があります。まず、アクション レジストリには先行投資が必要です。各アクションは、前提条件、効果、および責任の割り当てを指定して指定する必要があります。プロセスの文書化が不十分な組織の場合、この登録作業は多大な労力を要する可能性があります。当社では、既存のワークフローを監視し、人間によるレビューのためのアクション定義を提案する自動検出ツールを使用してこれを軽減します。第 2 に、12M パラメータの意図分類子は、トレーニング データで表現されていない新しいリクエスト タイプにうまく一般化できない可能性があります。私たちは、明確化プロトコルと、蓄積された生産データの定期的な再トレーニングを通じて、この問題に対処します。 3 番目に、座標ベースのシャーディング戦略は、クロスユニバース ルーティングがまれであることを前提としています。部門間のコラボレーションが頻繁に行われている組織では、3% の仮定が当てはまらない可能性があり、より洗練されたメタルーティング層が必要になります。
10. 結論
この文書では、MARIA OS に実装された Action Router の完全なエンジニアリング アーキテクチャについて説明しました。 3 層スタック (インテント パーサー、アクション リゾルバー、ゲート コントローラー) は、アクション ルーター インテリジェンス理論の理論的フレームワークを実稼働対応のシステムに変換します。各レイヤーには、明確に定義されたインターフェイス、レイテンシ バジェット、および独立したスケーラビリティがあります。
The recursive self-improvement mechanism demonstrates that the Action Router is not a static system but a learning one. The +4.4% accuracy improvement over 30 days, achieved through the principled application of online convex optimization with provable regret bounds, suggests that longer deployment periods will yield further gains, asymptotically approaching the optimal routing policy for the organization.
The scaling architecture proves that action-level routing is not inherently more expensive than keyword routing. Through coordinate-based sharding and hierarchical caching, the Action Router sustains 10,000 rps at sub-30ms P99 latency — performance that is 3.2× faster than semantic routing at the P99 and competitive with keyword routing, while delivering dramatically higher accuracy and responsibility compliance.
The integration with the Decision Pipeline state machine, formalized as a product automaton, ensures that every routing decision connects seamlessly to the governance infrastructure. No routed action bypasses the pipeline. No pipeline state is unreachable from the router. The routing layer and the governance layer are not separate systems bolted together; they are a single compositional architecture.
The Action Router is not the last word in intelligent routing. Future work includes multi-step planning (routing to action sequences rather than single actions), adversarial robustness (resistance to prompt injection attacks that attempt to manipulate routing), and federated learning across MARIA Galaxies (sharing routing knowledge across organizational boundaries without sharing private data). But the core insight — that routing must control actions, not classify words — is the foundation on which all future advances will build.
参考文献
1. さくら / 凡銀館 (2026) AI ルーティングがキーワード ベースではなくアクション ベースである必要がある理由。 note.com。 2. シャレフ・シュワルツ、S. (2012)。オンライン学習とオンライン凸最適化。 ML の基礎と傾向、4(2)、107-194。 3. アローラ、S. 他。 (2012年)。乗算重み更新メソッド。 コンピューティング理論、8(1)、121-164。 4. ホップクロフト、J.E. 他(2006)。 オートマトンの理論、言語、計算の紹介。第 3 版、ピアソン。 5. MARIA OS 技術文書 (2026)。アクションルーター実装ガイド、v1.0。 6. MARIA OS 技術文書 (2026)。デシジョン パイプライン ステート マシン仕様、v2.1。