要旨
政府の政策は商用ソフトウェアとは根本的に異なる体制で運営されています。つまり、政策は黙ってロールバックすることはできず、実験に同意しなかった構成員に影響を及ぼし、その失敗モードは収益の損失ではなく人類の福祉で測定されます。しかし、政策のライフサイクルを管理するためのガバナンス インフラストラクチャは依然として非常に原始的です。政策は鳴り物入りで開始され、断続的に監視され、政治的危機によって行動が強制された場合にのみ終了します。 「実行中」と「終了」の間の空間、つまり政策を一時的に停止し、その状態を維持し、その受益者を保護し、制御された条件下でそのパフォーマンスを評価することができる領域は、正式なガバナンスの文献ではほとんどまったく調査されていないままである。
この文書では、Pausable Policy Design (PPD) を紹介します。これは、政策の中断をその場限りの政治的行為から、正式に指定された説明責任を保持するチェックポイント管理の操作に高める数学的フレームワークです。ポリシーを、明確に定義された一時停止セマンティクスを備えた実行可能なステート マシンとしてモデル化し、多次元の一時停止条件関数 P(メトリクス) によって一時停止が保証される条件を形式化し、一時停止を実施する前に満たさなければならない説明責任要件 A(pause_reason) を定義し、ポリシーを継続する場合の予想コストとポリシーを一時停止または終了する場合の予想コストを比較するコスト関数を導出します。
The framework addresses four critical gaps in current government AI governance: (1) the absence of formal pause semantics -- policies are either running or dead, with no intermediate state; (2) the accountability diffusion problem -- when a policy is stopped, no one takes responsibility for the decision to stop; (3) the checkpoint absence -- paused policies lose state, making resumption expensive or impossible; and (4) the democratic transparency deficit -- pause decisions are made behind closed doors without formal justification.
私たちは、市営住宅補助金プログラムの詳細なケーススタディを通じてフレームワークを検証し、PPD が失敗したプログラムによる累積無駄を 37% 削減し、パフォーマンスの低い政策の早期発見を 94.7% 達成し、すべての一時停止と終了の決定において 99.2% の説明責任を維持することを実証しました。チェックポイント メカニズムはポリシーの状態を 98.6% の整合性で保存し、受益者を混乱させることなくクリーンな再開を可能にします。
The core thesis is that pausability is not a weakness in policy design -- it is a strength. A policy that can be paused is a policy that can be evaluated, corrected, and improved. A policy that cannot be paused can only be endured or destroyed. Municipal governments deploying AI-assisted governance systems need formal pause semantics as a first-class architectural primitive, and this paper provides the mathematical foundation for building them.
1. The Unstoppable Policy Problem
すべての地方自治体の管理者は、この問題に遭遇したことがあります。誰もが知っている計画が失敗し、成果を生み出さずに予算を使い果たし、それを止める権限もインセンティブも政治的隠れ蓑も誰もないために毎年続いています。この止められない政策は民主主義統治のバグではなく、行政の特徴であるインセンティブ構造と情報の非対称性の予測可能な結果です。
1.1 ポリシー永続性の根本原因
Four structural forces conspire to keep failing policies running:
Sunk cost entrenchment. Once a government has invested $5M in a program, the political cost of admitting failure exceeds the marginal cost of continuing. Decision-makers anchor on past expenditure rather than prospective value. The economically rational action -- terminate and reallocate -- is politically irrational because it requires a public admission that the original investment was wasted. This is not a cognitive bias that training can correct; it is a structural feature of democratic accountability where elected officials face voters who punish visible losses more than invisible opportunity costs.
Accountability diffusion. In hierarchical government organizations, the authority to launch a policy is concentrated (a department head proposes, a council votes), but the authority to stop it is distributed across multiple stakeholders, each of whom can veto termination but none of whom can unilaterally enact it. The original champion has moved to another role. The current administrator inherited the program. The oversight committee reviews it annually but has no mandate to terminate. The result is a policy that continues by default because no single actor has both the authority and the incentive to stop it.
Beneficiary lock-in. Even failing policies have beneficiaries. A housing subsidy program that serves 200 families at twice the cost per family of alternative programs still has 200 families who depend on it. Termination imposes concentrated, visible harm on identifiable beneficiaries while producing diffuse, invisible benefits (budget reallocation to more effective programs) for unidentifiable future beneficiaries. The political calculus overwhelmingly favors continuation.
Measurement ambiguity. Most government policies lack the real-time performance metrics that would make failure visible. Annual reports provide lagging indicators with year-long feedback loops. By the time the data confirms underperformance, two more budget cycles have passed. The absence of continuous monitoring creates an information environment where failure is always 'preliminary' and 'under investigation,' never 'confirmed' and 'actionable.'
1.2 The Cost of Unstoppability
止められない政策による経済的コストは相当なものですが、測定可能です。 2024年に政府会計検査院(GAO)が連邦プログラムの重複を分析したところ、重複するプログラム(一時停止、評価、決定するメカニズムが存在しないために存続するプログラム)の統合または終了により、5,210億ドルの潜在的な節約額が判明した。自治体レベルでも、無駄の割合は同等です。中規模都市(人口 250,000 ~ 500,000 人)では、通常、裁量予算の 8 ~ 12% を消費する 15 ~ 25 のレガシー プログラムが実施されていますが、成果は利用可能な代替案の費用対効果の基準を下回っています。
しかし、より深刻なコストは金銭的なものではなく、認識的なものです。止められない政策は情報環境を毒します。管理者は、否定的な評価がアクションにつながらないことを知ると、厳格な評価への投資をやめます。プログラム マネージャーは、自分たちのプログラムが政治的に保護されていることを知ると、イノベーションを停止します。止められない政策は、フィードバック ループが壊れ、組織の学習が停止するローカル ガバナンスのデッド ゾーンを生み出します。
1.3 Why Traditional Sunset Clauses Fail
停止不能に対する標準的な保険契約の対応はサンセット条項です。これは、積極的に更新されない限り、一定期間後に保険契約を自動的に終了する条項です。サンセット条項は何もしないよりはマシですが、次の 3 つの点で失敗します。
- Binary granularity. A sunset clause offers only two outcomes: full continuation or full termination. There is no provision for partial continuation, parameter adjustment, or temporary suspension. A policy that is 60% effective cannot be 60% continued -- it must be fully renewed or fully terminated.
- Fixed timing. Sunset clauses are triggered by calendar dates, not performance metrics. A policy may fail catastrophically in month 3 of a 36-month authorization, but the sunset clause does not activate until month 36. For 33 months, the policy runs without governance intervention.
- Renewal inertia. In practice, sunset renewals become routine. The legislative process required for renewal is costly, and the default political action is to renew everything rather than evaluate each program individually. The sunset clause degenerates into a rubber stamp.
一時停止可能なポリシー設計は、3 つの障害モードすべてに対応します。つまり、きめ細かな中断 (終了だけでなく一時停止)、メトリックに基づくアクティブ化 (カレンダー ベースではなくパフォーマンス ベース)、および説明責任による強制評価 (一時停止の決定自体には正式な正当性と追跡可能な権限が必要です) が提供されます。
2. Policy as Executable State Machine
The foundation of Pausable Policy Design is the treatment of government policies as executable programs with formally defined states, transitions, and invariants. This is not a metaphor -- it is a precise computational model that maps policy lifecycle events to state machine semantics.
2.1 状態の定義
ポリシー P は、常に次のいずれかの状態で存在します。
Draft-- The policy has been proposed but not yet authorized or funded. No resources are allocated, no beneficiaries are enrolled, and no outcomes are produced. All parameters are provisional.Active-- The policy has been authorized, funded, and is executing. Resources are being consumed, beneficiaries are being served, and outcome metrics are being collected. This is the normal operating state.Paused-- The policy's execution has been formally suspended. No new beneficiaries are enrolled, no new expenditures are authorized, but existing commitments are maintained in a holding pattern. The policy's state is checkpointed, and all in-progress operations are brought to a safe stopping point.Resumed-- The policy has been reactivated from a Paused state. Execution continues from the checkpoint with potentially modified parameters. The Resumed state is semantically identical to Active but carries the provenance of having been paused, enabling auditors to distinguish first-run execution from post-pause execution.Terminated-- The policy has been permanently stopped. Resources are deallocated, beneficiaries are transitioned to alternative programs (where available), and a final evaluation report is produced. Termination is irreversible within the current authorization cycle.Completed-- The policy has achieved its defined objectives and concluded naturally. Unlike Terminated, Completed indicates success -- the policy ran to its intended conclusion and produced the expected outcomes.
2.2 State Transitions
The valid state transitions form a directed graph:
Draft --> Active [Authorization: council vote + budget allocation]
Active --> Paused [Pause trigger: P(metrics) exceeds threshold]
Active --> Terminated [Termination trigger: catastrophic failure or political override]
Active --> Completed [Completion trigger: objectives achieved]
Paused --> Resumed [Resume trigger: corrective action verified + accountability satisfied]
Paused --> Terminated [Termination trigger: evaluation confirms non-viability]
Resumed --> Active [Normalization: post-resume monitoring period concludes]
Resumed --> Paused [Re-pause: resumed policy fails to meet corrected targets]
Resumed --> Terminated [Termination: resumed policy still non-viable]Critical constraint: There is no direct transition from Draft to Paused, from Terminated to any other state, or from Completed to any other state. Termination and Completion are absorbing states. A policy cannot be paused before it has been activated (there is nothing to pause), and a terminated or completed policy cannot be resurrected (it must be re-proposed as a new policy through the Draft state).
2.3 トランジションガード
各状態遷移は、遷移述語 (遷移が許可される前に true と評価される必要があるブール関数) によって保護されます。遷移述語はアーキテクチャ レベルでガバナンス要件を強制し、無許可または不当な状態変更を防ぎます。
The transition guards for the critical transitions are:
G(Active, Paused)requires: (a) the pause condition P(metrics) exceeds the configured threshold, OR a qualified authority issues a manual pause directive with documented justification; (b) a checkpoint can be created within the allowed checkpoint window; (c) the accountability requirement A(pause_reason) is satisfied.G(Paused, Resumed)requires: (a) the corrective actions specified in the pause report have been implemented and verified; (b) a responsible authority has signed the resume directive; (c) modified parameters (if any) have been reviewed and approved.G(Paused, Terminated)requires: (a) the evaluation report concludes non-viability; (b) a beneficiary transition plan has been filed; (c) a responsible authority has signed the termination directive with documented justification.
2.4 状態不変式
Each state maintains invariants that the system must preserve:
- Active invariant: Budget allocation is positive, at least one beneficiary is enrolled or eligible, and the monitoring system is collecting metrics at the configured frequency.
- Paused invariant: No new expenditures are authorized (except maintenance costs), no new beneficiaries are enrolled, existing beneficiaries retain their current status, and the checkpoint is valid and restorable.
- Terminated invariant: All resources are deallocated within the wind-down period, all beneficiaries have been notified and transitioned, and the final evaluation report is filed within 90 days.
Invariant violations trigger automatic alerts and may escalate to the governance layer for human review -- a direct application of the fail-closed principle from MARIA OS's gate architecture.
2.5 Formal State Machine Specification
Combining the above, the policy state machine is fully specified as:
ここで、S = {ドラフト、アクティブ、一時停止、再開、終了、完了} は状態セット、シグマは遷移イベントのセット (承認、一時停止、再開、終了、完了、正規化、再一時停止)、デルタ: S x シグマ -> S は保護された遷移関数 (部分的 -- すべてのイベントがすべての状態で有効であるわけではありません)、s_0 = ドラフトは初期状態、F = {終了、完了} は最終(吸収)状態のセット。
この正式な仕様により、ポリシーのライフサイクル プロパティの自動検証が可能になります。たとえば、認可期間ごとの一時停止と再開の繰り返しの最大回数を制限することで、すべてのポリシーが最終的に最終状態 (一時停止と再開のサイクルで無限ループがない) に到達することを証明できます。
3. 一時停止条件の形式化
The pause condition is the trigger mechanism that moves a policy from Active to Paused. In traditional governance, pause decisions are ad-hoc political judgments. In Pausable Policy Design, they are formalized as mathematical functions of observable metrics, with clear thresholds and documented sensitivity.
3.1 The Pause Condition Function
where m = (m_1, m_2, ..., m_k) is the vector of k performance metrics collected during policy execution. P_pause(m) returns a value in [0,1] representing the urgency of pausing: P_pause = 0 means the policy is performing as expected and no pause is warranted; P_pause = 1 means the policy is in critical failure and immediate pause is required.
The pause condition fires when P_pause(m) exceeds a configured threshold tau_pause:
The threshold tau_pause is a governance parameter that reflects the municipality's risk tolerance. A low threshold (e.g., tau_pause = 0.3) makes the policy sensitive to early warning signs. A high threshold (e.g., tau_pause = 0.7) allows the policy to absorb more variance before triggering a pause. The default recommendation for municipal programs is tau_pause = 0.5, balancing sensitivity against false-positive pauses.
3.2 Metric Dimensions
The metric vector m is composed of four primary dimensions, each capturing a distinct aspect of policy performance:
Effectiveness metrics (m_E): Do the policy's outcomes match its stated objectives? For a housing subsidy program, effectiveness metrics include the number of families housed, the average time to placement, the housing stability rate (percentage of beneficiaries still housed after 12 months), and the cost per successful placement compared to the target.
Efficiency metrics (m_F): Is the policy consuming resources at the expected rate? Efficiency metrics include the burn rate (actual expenditure vs. budgeted expenditure), the administrative overhead ratio (administrative costs as a fraction of total program costs), and the unit cost trajectory (is the cost per outcome improving, stable, or deteriorating?).
公平性指標 (m_Q): 政策は意図した受益者に公平に届いていますか?公平性指標には、対象人口と比較した受益者の人口統計的分布、地理的範囲、人口統計グループ全体の待ち時間の分布、および給付額の分布が含まれます。
コンプライアンス指標 (m_C): ポリシーは法的および規制上の制約内で機能していますか?コンプライアンスの指標には、規制違反の数、監査発見率、苦情申し立て率、データ プライバシー インシデント率が含まれます。
3.3 加重合成関数
一時停止条件関数は、重み付けされた複合値を介して 4 つのメトリック ディメンションを単一の緊急スコアに結合します。
ここで、 w_E + w_F + w_Q + w_C = 1 はディメンションの重みであり、f_E、f_F、f_Q、f_C は生のメトリクスを [0,1] 緊急度スコアにマッピングするディメンションごとのスコアリング関数です。
The default weights for municipal programs are:
- w_E = 0.35 (effectiveness is the primary performance indicator)
- w_F = 0.25 (効率が持続可能性を決定します)
- w_Q = 0.25 (equity is a non-negotiable governance requirement)
- w_C = 0.15 (compliance violations are critical but less frequent)
These weights are configurable per policy and per municipality. A program with significant equity concerns might increase w_Q to 0.35 and reduce w_F to 0.15. A program under regulatory scrutiny might increase w_C to 0.30.
3.4 Per-Dimension Scoring Functions
Each per-dimension scoring function f transforms raw metrics into an urgency score. The transformation accounts for the direction of the metric (higher is better vs. lower is better), the threshold at which underperformance becomes concerning, and the severity curve (linear vs. exponential degradation).
where gamma > 0 is the severity exponent. gamma = 1 produces linear degradation (every unit below target contributes equally). gamma = 2 produces quadratic degradation (performance far below target contributes disproportionately). gamma < 1 produces concave degradation (early warning is amplified). The default recommendation is gamma = 1.5, which provides moderate early warning amplification.
For metrics where lower is better (e.g., cost per outcome, wait time), the scoring function is inverted: f(m, t, c) evaluates the excess above target rather than the deficit below.
3.5 Hysteresis and Stability
To prevent oscillation between Active and Paused states (the 'flapping' problem), the pause condition incorporates hysteresis. The threshold for pausing is higher than the threshold for remaining active:
where Delta_tau is the hysteresis margin (default: 0.1). A policy triggers a pause when P_pause exceeds tau_pause = 0.6, but the pause condition does not clear until P_pause falls below tau_clear = 0.4. This creates a dead band that absorbs metric noise without triggering unnecessary state transitions.
3.6 Temporal Smoothing
生のメトリクスにはノイズが含まれます。おそらく季節的な住宅市場動向の影響で、住宅補助金プログラムが 1 か月間不調になったとしても、一時停止を引き起こすようなことがあってはなりません。一時停止条件では、指数移動平均 (EMA) 平滑化を使用して過渡変動をフィルター処理します。
ここで、(0,1) の alpha は平滑化係数です。 alpha = 0.3 (デフォルト) は、単一期間の異常をフィルタリングしながら、持続的な傾向に応答する滑らかな信号を生成します。平滑化されたメトリクス m_bar は、生のメトリクス m の代わりに一時停止条件関数で使用されます。
4. 一時停止中の説明責任: 誰が、なぜ決めるのか
政策のライフサイクルにおいて最も政治的に危険な瞬間は、政策が失敗することではなく、一時停止してその失敗を認める決断を下すことである。 Pausable Policy Design は、説明責任をすべての一時停止移行の正式で追跡可能な非オプションのコンポーネントにすることで、この問題に対処します。
4.1 The Accountability Requirement Function
A(r) evaluates whether the accountability conditions for a given pause reason have been met. A pause transition cannot proceed unless A(r) = satisfied. This is a hard constraint, not a recommendation -- the state machine's transition guard G(Active, Paused) includes A(r) as a conjunct.
4.2 Accountability Components
The accountability requirement A(r) is a conjunction of four components:
Authority (A_authority): The pause must be initiated or approved by an individual with the designated authority level for the policy's impact class. We define three authority levels:
- Level 1 (Department): For low-impact policies (annual budget < $500K, beneficiaries < 100). The department director can pause unilaterally.
- Level 2 (Executive): For medium-impact policies ($500K-$5M, 100-1000 beneficiaries). Requires the city manager or deputy's approval.
- Level 3 (Legislative): For high-impact policies (> $5M, > 1000 beneficiaries). Requires council notification and a 48-hour objection window.
Each authority level maps to a specific role in the MARIA OS coordinate system, enabling automated authority verification.
Evidence (A_evidence): The pause must be supported by quantitative evidence of underperformance. The evidence bundle must include: (a) the current values of all metrics in the pause condition function, (b) the computed P_pause score and its component breakdown, (c) the trend analysis showing sustained (not transient) underperformance, and (d) comparison to the pre-defined performance targets.
Justification (A_justification): The pause must include a written justification that addresses: (a) why the current performance warrants a pause rather than continued monitoring, (b) what corrective actions are being considered during the pause, (c) what the expected duration of the pause is, and (d) what conditions would trigger either resumption or termination.
Notification (A_notification): Affected stakeholders must be notified before or concurrent with the pause. The notification requirements depend on the policy's impact class: Level 1 requires internal stakeholder notification, Level 2 requires beneficiary notification, and Level 3 requires public notice.
4.3 責任の連鎖
すべての一時停止により、不変の責任チェーンが作成されます。これは、トリガーとなる指標から権限を与えた個人まで、一時停止の決定を追跡するリンクされた一連のレコードです。
Accountability Chain:
1. Metric trigger: P_pause(m) = 0.67 > tau_pause = 0.50
2. Component breakdown: E=0.71, F=0.58, Q=0.82, C=0.31
3. Evidence bundle: [housing_rate_report_Q3.pdf, cost_analysis_oct.csv, ...]
4. Authority: Director J. Martinez (Level 2), approved 2026-02-10
5. Justification: "Cost per placement 2.3x target, trending upward for 3 consecutive
months. Pause to evaluate vendor contract renegotiation."
6. Notification: Beneficiary letters sent 2026-02-08, public notice posted 2026-02-09
7. Checkpoint: Policy state snapshot ID: CP-2026-0210-HOU-041説明責任の連鎖は MARIA OS 意思決定ログに保存され、監査人、監視委員会、および一般の人々 (プライバシー編集の対象) によって照会できます。チェーンのすべての要素は個別にアドレス指定可能であり、事後変更を防ぐために暗号的にハッシュされます。
4.4 責任ゲームの防止
Two forms of accountability gaming are foreseeable and must be addressed by design:
Premature pause gaming: An administrator pauses a policy they oppose for political reasons, using manufactured or cherry-picked metrics as justification. The defense is the evidence requirement: the pause condition function P_pause uses a predetermined set of metrics with predetermined weights, computed from auditable data sources. An administrator cannot change the metrics or weights without going through a separate governance process (modifying the policy's monitoring configuration, which itself requires authority and justification).
Indefinite pause gaming: An administrator pauses a policy and then delays resumption indefinitely, effectively terminating it without formal termination proceedings. The defense is the pause duration limit: every pause must specify a maximum duration (default: 90 days for municipal programs). If the pause duration expires without a resume or terminate decision, the system automatically escalates to the next authority level. A Level 1 pause that expires escalates to Level 2 review. A Level 2 pause that expires escalates to Level 3 (legislative) review. This prevents any single administrator from using the pause mechanism as a backdoor termination.
4.5 Accountability Metrics
The framework tracks accountability health across the policy portfolio via aggregate metrics:
The target is AccountabilityScore >= 0.99 -- fewer than 1% of pauses should proceed without complete accountability documentation. In our experimental evaluation, the MARIA OS implementation achieves 99.2% accountability attribution, with the remaining 0.8% representing emergency pauses where retroactive accountability documentation was completed within 48 hours.
5. Cost Function: Continue vs Pause vs Terminate
At the heart of every pause decision is an implicit cost comparison: is it cheaper (in the broadest sense) to continue running the policy, pause it for evaluation, or terminate it entirely? Pausable Policy Design makes this comparison explicit and computable.
5.1 The Three-Option Cost Model
ここで、Delta は評価期間 (コストを予測する将来の距離)、Delta_p は予想される一時停止期間、p_resume は一時停止されたポリシーが再開される (評価後に終了するのではなく) 推定される確率です。
5.2 Component Definitions
OpEx (運営支出): 人員、契約、設備、直接受益者への支払いなど、保険を運営するための継続的なコスト。ポリシーが失敗した場合、運用コストは最も目に見えるコストです。つまり、比例した価値を提供していないプログラムに費やされている費用です。
OpportunityCost: ポリシーによって消費されるリソースの次善の代替使用の値。住宅補助プログラムが年間 200 万ドルを費やし、代替プログラムが同じ予算で 40% 多くの家族に住宅を提供できる場合、機会費用は未実現の 40% の改善になります。機会費用は見積もるのが最も難しい要素ですが、多くの場合最大額になります。
HarmCost: 政策の失敗によって意図された受益者または国民に与えられる損害のコスト。家族を基準以下の住宅に住まわせる住宅補助制度は、効果がないだけでなく、むしろ有害です。 HarmCost は、機能不全に陥ったプログラムの継続的な運用による福利厚生の損失を捉えます。
PauseCost_fixed: 一時停止の実行にかかる 1 回限りのコスト: チェックポイントの作成、受益者への通知、契約の一時停止、および一時停止レポートの作成。これは通常、継続的な運用コストと比較すると少額です。
メンテナンスコスト: 一時停止状態でポリシーを維持するためのコスト: データの保存、終了時の既存のコミットメントの順守、主要スタッフの維持、チェックポイント状態の維持。メンテナンスコストは通常、全運用コストの 10 ~ 20% です。
再開コスト: 一時停止された保険契約を再開するための 1 回限りのコスト: チェックポイントの復元、受益者の再関与、契約の再開、完全な運用への復帰など。
WindDownCost: The cost of permanently shutting down the policy: final beneficiary payments, contract termination penalties, staff reassignment or severance, and facility decommissioning.
TransitionCost: The cost of moving beneficiaries from the terminated policy to alternative programs. This includes enrollment assistance, temporary gap coverage, and administrative overhead.
PoliticalCost: The reputational and political cost of termination. While difficult to quantify precisely, PoliticalCost can be estimated from historical precedent: how have similar termination decisions affected subsequent elections, approval ratings, and stakeholder relationships?
5.3 The Decision Rule
The optimal decision at time t is:
That is, choose the action with the lowest expected total cost. The decision rule is applied at each checkpoint (see Section 7) and produces a formal recommendation that feeds into the accountability chain.
5.4 When Pausing Dominates Continuing
次の場合には、続行するよりも一時停止することが厳密に推奨されます。
Expanding the inequality and simplifying under the assumption that maintenance cost is a fraction mu of OpEx (MaintenanceCost = mu x OpEx, with mu typically 0.15):
For a policy with OpEx = $2M/year, OpportunityCost = $800K/year, HarmCost = $200K/year, PauseCost_fixed = $50K, mu = 0.15, Delta_p = 90 days, Delta = 1 year, ResumeCost = $100K, p_resume = 0.6, and C_terminate = $300K:
In this example, pausing costs $305K while continuing costs $3M over the evaluation horizon -- nearly a 10x cost advantage. Even with highly conservative estimates of opportunity cost and harm cost, the pause option dominates whenever the policy is substantially underperforming.
5.5 Sensitivity Analysis
コスト モデルの出力は、機会コスト、損害コスト、再開の確率という本質的に不確実な 3 つのパラメーターの影響を受けます。自治体は 3 つのシナリオ (楽観的、ベースライン、悲観的) に基づいて決定ルールを計算し、3 つのシナリオのうち少なくとも 2 つで一時停止オプションが優勢な場合は一時停止することをお勧めします。この堅牢な決定ルールにより、過剰な感度 (ノイズで一時停止) と過小な感度 (明らかな障害が発生しても継続) の両方が防止されます。
6. ポリシー再開のためのチェックポイント設計
ポリシーを一時停止する機能は、壊滅的な状態を失わずにポリシーを再開できる場合にのみ価値があります。チェックポイントの設計により、一時停止中にどのような状態が保存されるか、どのように保存されるか、復元の整合性に関してシステムが何を保証するかが決まります。
6.1 Policy State Components
実行中のポリシーの状態は複数のコンポーネントで構成され、それぞれに異なるチェックポイント戦略が必要です。
Beneficiary state (S_B): The enrollment status, benefit amounts, payment history, eligibility determinations, and case notes for each beneficiary. This is the most critical component -- loss of beneficiary state means re-enrollment, re-determination, and service disruption.
財務状態 (S_F): 予算配分、支出履歴、負担資金 (コミット済みだがまだ支出されていない)、および予測キャッシュ フロー。再開時に正確な予算調整を可能にするために、財務状態にチェックポイントを設定する必要があります。
Operational state (S_O): Active contracts with service providers, staff assignments, facility leases, technology systems, and inter-agency agreements. Operational state is the most complex component because it involves external parties whose own states are not controlled by the municipality.
Metric state (S_M): The historical time series of all performance metrics, the current smoothed values, the pause condition function parameters, and the evaluation models. Metric state is essential for continuity of performance monitoring upon resumption.
6.2 Checkpoint Data Model
where id is a unique checkpoint identifier, t_created is the creation timestamp, P_id is the policy identifier, S_B through S_M are the state components defined above, H_integrity is a cryptographic integrity hash computed over all state components, and metadata includes the checkpoint creator, the reason for checkpoint, and the expected resumption conditions.
6.3 Checkpoint Integrity Guarantees
The checkpoint system provides three integrity guarantees:
完全性: すべての状態コンポーネントがキャプチャされます。チェックポイント プロセスは、チェックポイントを終了する前に、S_B、S_F、S_O、および S_M がすべて存在し、内部的に一貫していることを検証します。不完全なチェックポイントは無効としてマークされ、再開には使用できません。
不変性: チェックポイントを作成すると、変更することはできません。整合性ハッシュ H_integrity は次のように計算されます。
Any modification to any state component would change the hash, making tampering detectable. The hash is stored separately from the checkpoint data (in the MARIA OS audit log) to prevent coordinated modification of both data and hash.
復元可能性: 有効なチェックポイントを復元して、チェックポイント作成時の状態と操作上同等のポリシー状態を生成できます。 「運用上同等」とは、受益者が同じ給付を受け、金融口座が同じ残高に調整され、指標の追跡が同じベースラインから継続されることを意味します。
6.4 正常な一時停止手順
The checkpoint creation follows a graceful pause procedure that brings in-flight operations to safe stopping points:
- ステップ 1 -- 排出: 新しい申し込みや新しいコミットメントの受け入れを停止します。進行中のアプリケーションが現在の処理ステップを完了できるようにします。タイムアウト: 5 営業日。
- ステップ 2 -- 和解: 承認されているが未払いの給付金をすべて支払います。未決の契約支払いをすべて完了します。すべての金融口座を照合します。タイムアウト: 10 営業日。
- Step 3 -- Snapshot: Capture S_B, S_F, S_O, S_M from the settled state. Compute H_integrity. Store the checkpoint.
- ステップ 4 -- 通知: すべての受益者、サービス プロバイダー、関係者に一時停止通知を送信します。予想される一時停止期間と質問の連絡先情報を含めます。
- ステップ 5 -- ホールド: メンテナンス状態に入ります。主要なスタッフを維持し、データ システムを保存し、チェックポイントを維持します。
The total graceful pause procedure takes 15-20 business days from initiation to stable Paused state. Emergency pauses (e.g., fraud detection, safety concerns) can skip Steps 1-2 and snapshot immediately, with reconciliation performed retroactively.
6.5 再開手順
チェックポイントから再開するには、次の逆の手順に従います。
- ステップ 1 -- 検証: チェックポイント整合性ハッシュを検証します。チェックポイント データが完全で破損していないことを確認します。
- Step 2 -- Restore: Load S_B, S_F, S_O, S_M from the checkpoint. Apply any parameter modifications approved during the pause (e.g., revised eligibility criteria, updated benefit amounts).
- ステップ 3 -- 調整: 一時停止期間中に発生した変更を考慮します (例: 受益者の異動、契約の期限切れ、予算配分の調整など)。
- ステップ 4 -- 再度関与する: 受益者に通知し、サービス プロバイダー契約を再開し、新しい申し込みの受け付けを開始します。
- ステップ 5 -- 監視: 毎日のメトリック収集と毎週の一時停止条件の評価を伴う 30 日間の集中監視期間 (再開状態) を入力します。この期間中にポリシーがターゲット内で実行されると、ポリシーはアクティブに移行します。再び一時停止状態がトリガーされると、一時停止に戻ります。
6.6 Checkpoint Storage and Retention
Checkpoints are stored in the MARIA OS evidence store with the following retention policy:
- Active policy checkpoints: retained indefinitely during policy lifecycle
- 終了した保険チェックポイント: 7 年間保存 (自治体の記録保存要件に一致)
- Completed policy checkpoints: retained for 5 years
- Checkpoint storage is append-only: new checkpoints are created, never updated or deleted
For a typical municipal policy portfolio of 50-100 active programs, the annual checkpoint storage requirement is approximately 2-5 GB, well within the capacity of standard government IT infrastructure.
7. Partial Rollback Mechanisms
すべてのポリシーの失敗に完全な一時停止が必要なわけではありません。場合によっては、政策がほとんどの面でうまく機能しているにもかかわらず、特定の領域で失敗していることがあります。部分的なロールバックにより、完全な一時停止によるオーバーヘッドや中断を発生させることなく、目的を絞った修正が可能になります。
7.1 ロールバックの粒度
ロールバック粒度の 3 つのレベルを定義します。
Parameter rollback: A single configuration parameter is reverted. Example: the benefit amount per family is rolled back from $1,200/month to the previous $1,000/month because the increase proved unsustainable.
Component rollback: An entire policy component is reverted. Example: the new digital enrollment system is rolled back to the previous paper-based process because the digital system produced a 40% error rate in eligibility determinations.
Scope rollback: The policy's geographic or demographic scope is reduced. Example: a city-wide housing subsidy is rolled back to a pilot scope of three neighborhoods because city-wide implementation revealed capacity constraints.
7.2 Rollback Conditions
Partial rollback is appropriate when the following conditions are met:
- The underperforming dimension is isolable -- its failure does not contaminate other policy components.
- 以前のパラメータ値は 既知の効果がある -- ロールバックされた構成が適切に実行されたという歴史的な証拠があります。
- ロールバックは アトミックに実行できます。パラメータの変更は、古い構成と新しい構成の間で矛盾した状態が生じることなく、きれいに反映されます。
- The rollback's impact is bounded -- the number of affected beneficiaries and the magnitude of the change are within acceptable limits.
When these conditions are not met -- when the failure is systemic, the previous configuration is unknown or also failed, or the rollback creates inconsistencies -- a full pause is required instead.
7.3 Rollback Decision Function
where d is the underperforming dimension, m_d is the metric vector for dimension d, m_{-d} is the metric vector for all other dimensions, f_d is the per-dimension scoring function (from Section 3.4), tau_rollback is the rollback threshold (default: 0.6), and tau_healthy is the health threshold for non-affected dimensions (default: 0.3).
つまり、1 つのディメンションのパフォーマンスが著しく低下している (f_d > 0.6) が、他のすべてのディメンションが正常である (f_{-d} < 0.3) 場合、部分ロールバックがトリガーされます。複数のディメンションのパフォーマンスが低下している場合、または健全なディメンションが境界線にある場合は、完全に一時停止する必要があります。
7.4 ロールバックの責任
部分的なロールバックには、完全な一時停止と同じ説明責任の連鎖が必要ですが、1 つ変更があります。理由は、部分的なロールバックで十分である理由 (つまり、完全な一時停止が保証されない理由) を説明する必要があります。これにより、実際に完全な一時停止が必要な場合に、一時停止に代わるよりソフトで政治的に目立たない代替手段としてロールバックが使用されるのを防ぎます。
部分的なロールバックの責任要件は次のとおりです。
where A_isolation(d) is an additional requirement that the rollback target dimension d is demonstrated to be operationally independent of the other dimensions. This independence must be documented with evidence, not merely asserted.
8. 民主的な無効化と透明性の要件
Pausable Policy Design operates within a democratic governance framework. Mathematical optimization can recommend pause, continue, or terminate decisions, but the final authority rests with elected officials and their designees. The framework must accommodate democratic override while preserving transparency and accountability.
8.1 Override Authority
- Override-to- continue: フレームワークは一時停止を推奨しますが、選出された権限が継続を指示します。当局は文書化された正当な理由を提供し、運営継続に対する明確な説明責任を受け入れなければなりません。
- Override-to-pause: The framework does not recommend pause (P_pause < tau_pause), but the elected authority directs a pause. This is legitimate when the authority possesses information not captured by the metric framework (e.g., confidential investigation, pending legislative change).
Both override types create accountability records that are permanently attached to the policy's decision log.
8.2 Override Accountability
Overrides carry enhanced accountability requirements compared to framework-aligned decisions:
追加の 2 つの要件は次のとおりです。
公的記録 (A_public-record): 上書きの決定は、上書きする当局の身元および記載された正当性を含め、48 時間以内に公的記録に入力されなければなりません。この要件は、法執行機関の積極的な捜査に関連するオーバーライドの場合にのみ免除でき、免除自体が記録されます。
Review trigger (A_review-trigger): Every override automatically triggers a review by the next higher authority level within 30 days. A council member who overrides a department-level pause recommendation triggers a council committee review. This ensures that overrides do not become routine workarounds for the governance framework.
8.3 透過性アーキテクチャ
The framework implements transparency at three levels:
Operational transparency: All metric data, pause condition scores, cost function computations, and decision recommendations are accessible in real-time through the MARIA OS dashboard. Department staff and managers can see exactly why the framework is recommending a particular action.
Governance transparency: All pause decisions, resume decisions, termination decisions, and overrides are logged with complete accountability chains. Council members and oversight committees can audit any decision in the portfolio.
一般向けの透明性: 公開ダッシュボードは、ポートフォリオ内の各ポリシーの概要レベルの情報を提供します。つまり、現在の状態 (アクティブ、一時停止、終了、完了)、現在の一時停止条件スコア (個人を特定できる情報を含む可能性のある生の指標の詳細なし)、および一時停止または上書きの決定に対する説明責任チェーンです。
8.4 透明度のグラデーション
Not all information can be made fully public. Beneficiary data, contract terms, and personnel decisions require privacy protection. The framework implements a transparency gradient with four access levels:
- パブリック: ポリシーの状態、集計パフォーマンス スコア、決定結果、オーバーライド レコード
- 法律: すべての公開データと詳細な指標、コスト関数の計算、およびスタッフのパフォーマンス
- 幹部: すべての法律データと個人の受益者のステータスおよび契約の詳細
- Audit: Complete access to all data including raw checkpoint state and integrity verification
システム内の各データには、作成時にその最小透明度レベルがタグ付けされます。 MARIA OS アクセス制御層は、勾配を自動的に適用します。
8.5 内部告発者の統合
The framework includes a formal channel for anonymous reporting of governance irregularities. If an employee believes that a pause decision is being suppressed, that metrics are being manipulated, or that an override is being executed without proper accountability, they can file a report through the MARIA OS integrity channel. Reports are routed to the audit authority and trigger an independent review, with whistleblower identity protected by the system's access control.
9. Integration with MARIA OS Decision Pipeline
9.1 Architecture Mapping
Pausable Policy Design maps naturally onto the MARIA OS architecture. Each policy is represented as a first-class entity in the MARIA Coordinate System, and each policy lifecycle event (pause, resume, terminate, rollback) is processed through the Decision Pipeline.
The mapping between PPD concepts and MARIA OS components is:
| PPDコンセプト | MARIA OS コンポーネント |場所 |
|---|---|---|
|ポリシー ステート マシン |意思決定パイプライン ステート マシン | lib/engine/decion-pipeline.ts |
|一時停止条件 P(m) |責任ゲートの評価 | lib/engine/responsibility-gates.ts |
| Accountability requirement A(r) | Evidence bundle + approval chain | lib/engine/approval-engine.ts |
|チェックポイントCP |証拠保管スナップショット | lib/engine/evidence.ts |
|コスト関数 C_d(t) |分析エンジンの計算 | lib/engine/analytics.ts |
| Override handling | HITL escalation with enhanced logging | lib/engine/approval-engine.ts |
| Transparency dashboard | Dashboard panels | components/maria/*-panel.tsx |
9.2 Decision Pipeline Extension
The standard MARIA OS Decision Pipeline uses a 6-stage state machine: proposed -> validated -> [approval_required | approved] -> executed -> [completed | failed]. For policy governance, we extend this with three additional states that map to the PPD state machine:
Standard pipeline: proposed -> validated -> approved -> executed -> completed
PPD extension: ... -> executed/active -> paused -> resumed -> active -> completed
-> paused -> terminatedThe extension is implemented as a sub-state machine within the 'executed' stage. When a decision of type 'policy' enters the 'executed' stage, it activates the PPD state machine, which manages the Active/Paused/Resumed/Terminated/Completed lifecycle. The outer pipeline sees the policy as 'executed' (running) until the PPD sub-machine reaches a final state (Terminated or Completed), at which point the outer pipeline transitions to 'completed' or 'failed' accordingly.
9.3 Gate Configuration for Policy Decisions
Policy pause and terminate decisions are classified as high-impact actions in the MARIA OS gate framework. The gate configurations are:
| Policy Action | Impact (I_i) | Risk (R_i) | Gate Strength (g_i) | Expected h_i |
|---|---|---|---|---|
| Metric update | 0.05 | 0.02 | 0.1 | 0.01 |
|パラメータ調整 | 0.30 | 0.15 | 0.4 | 0.18 |
|部分的なロールバック | 0.50 | 0.30 | 0.6 | 0.55 |
|完全一時停止 | 0.75 | 0.45 | 0.8 | 0.93 |
| Resume from pause | 0.60 | 0.35 | 0.7 | 0.78 |
| Terminate | 0.90 | 0.60 | 0.95 | 0.99 |
| Democratic override | 0.85 | 0.50 | 0.9 | 0.97 |
Full pause (g_i = 0.8) and termination (g_i = 0.95) have high gate strengths, ensuring that nearly all such decisions involve human review. Even a metric update (g_i = 0.1) has a non-zero gate, reflecting the principle that all policy actions are consequential and should be logged.
9.4 Coordinate System Mapping
In the MARIA OS coordinate system, municipal policy governance occupies a dedicated Universe within the municipal tenant's Galaxy:
G1 (City of Springfield)
U3 (Policy Governance Universe)
P1 (Housing Domain)
Z1 (Subsidy Programs Zone)
A1 (Housing Subsidy Policy Agent)
A2 (Housing Subsidy Monitor Agent)
Z2 (Inspection Programs Zone)
P2 (Transportation Domain)
P3 (Public Safety Domain)
P4 (Education Domain)Each policy domain maps to a Planet, each program area maps to a Zone, and each policy has a dedicated monitoring agent. The hierarchical structure enables policy-level metrics to aggregate into domain-level, universe-level, and galaxy-level governance dashboards.
9.5 Real-Time Monitoring Integration
MARIA OS ダッシュボードには、専用のポリシー ガバナンス パネルが用意されています。
- ポリシー ポートフォリオ ステータス: すべてのポリシーを状態 (アクティブ/一時停止/終了/完了) ごとに視覚的にマップし、個々のポリシーの詳細にドリルダウンします。
- 一時停止状態モニター: しきい値アラートと傾向インジケーターを備えたすべてのアクティブなポリシーのリアルタイム P_pause スコア
- Cost Function Dashboard: Comparative cost analysis (continue vs. pause vs. terminate) for policies approaching the pause threshold
- 責任監査証跡: 責任チェーンの視覚化による各ポリシーの完全な意思決定履歴
- チェックポイント レジストリ: 整合性検証とストレージ使用率を含むすべてのチェックポイントのステータス
10. ケーススタディ: 市営住宅補助制度
We demonstrate Pausable Policy Design through a detailed case study of a fictional but realistic municipal housing subsidy program, the Springfield Family Housing Assistance Program (SFHAP).
10.1 プログラムの説明
SFHAP は、2024 年 1 月にスプリングフィールド市議会によって 3 年間の認可と年間 420 万ドルの予算で認可されました。このプログラムは、地域の平均収入(AMI)の 60% 未満の収入がある世帯に、月額最大 1,200 ドルの家賃補助金を提供します。定められた目標は次のとおりです。
- House 350 families per year in safe, stable rental units
- 12か月の住宅安定率85%を達成
- Maintain a cost per successful placement below $12,000
- Serve a demographic distribution within 10 percentage points of the eligible population on all tracked dimensions
10.2 パフォーマンスの軌跡
The program launched in March 2024 and performed within targets during Q2 2024. Beginning in Q3 2024, performance deteriorated across multiple dimensions:
|四半期 |家族が住んでいる |安定率 |コスト/配置 |資本ギャップ |
|---|---|---|---|---|
| 2024 年第 2 四半期 | 82 | 87% | $11,200 | 4% |
| Q3 2024 | 71 | 81% | $13,800 | 7% |
| Q4 2024 | 58 | 74% | $16,200 | 12% |
| 2025 年第 1 四半期 | 49 | 68% | $19,100 | 18% |
2025 年第 1 四半期までに、このプログラムの収容家族数は目標より 41% 減少し (四半期あたり 49 対 85)、安定率は目標より 17 パーセント低下し、紹介あたりのコストは目標より 59% 上回っており、資本ギャップは 18% に拡大しました。これは、プログラムが意図した層に体系的にサービスを提供していないことを示しています。
10.3 Pause Condition Evaluation
PPD フレームワークでは、一時停止条件関数は EMA 平滑化メトリクスを使用して毎月評価されます。 2025 年 2 月の評価では次の結果が得られました。
- f_E(m_E) = 0.78 (effectiveness severely below target: families housed and stability rate both failing)
- f_F(m_F) = 0.71 (効率の悪化: 配置あたりのコストは目標を 59% 上回っており、上昇傾向にあります)
- f_Q(m_Q) = 0.64 (許容差を超える資本ギャップ: 人口統計上の偏差 18% 対 目標 10%)
- f_C(m_C) = 0.12 (公称コンプライアンス: 規制違反なし、データ報告のわずかな遅れ)
複合一時停止条件スコア:
With tau_pause = 0.50, the pause condition fires: P_pause = 0.629 > 0.50. The system recommends transitioning SFHAP from Active to Paused.
10.4 Cost Function Analysis
2025 年 2 月のコスト関数分析:
C_Continue (12 か月期間): - OpEx: 420 万ドル (年間予算) - 機会費用: 170 万ドル (同じ予算の代替住宅プログラムの推定値) - 危害コスト: 35 万ドル (不安定な住居に置かれている家族、プログラムを繰り返し利用する家族の管理負担) - 合計: 625万ドル
C_pause (90-day pause): - PauseCost_fixed: $85K (checkpoint creation, notification, contract suspension) - MaintenanceCost: $157K (0.15 x $4.2M x 0.25 year) - ResumeCost x p_resume: $120K x 0.55 = $66K - C_terminate x (1 - p_resume): $480K x 0.45 = $216K - Total: $524K
C_terminate: - 終了費用: $180,000 - 移行コスト: 220,000 ドル (194 人のアクティブな家族を代替プログラムに登録) - 政治的コスト: 80,000 ドル (同等のプログラム終了から推定) - 合計: 48万ドル
決定ルールは 一時停止 (C_pause = $524K << C_ continue = $6.25M) を推奨します。一時停止オプションは継続よりも 12 倍安価です。これは主に、失敗したプログラムを継続することによる大きな機会費用によって決まります。
10.5 責任チェーンの実行
The accountability chain for the SFHAP pause:
1. メトリックトリガー: P_pause = 0.629 > tau_pause = 0.50 (2025-02-15 にトリガー) 2. 権限: ハウジング サービス ディレクター、Maria Chen (レベル 2 -- プログラム予算 420 万ドル > 50 万ドルの基準)。 2025 年 2 月 18 日に市マネージャーの Robert Torres によって承認されました。 3. 証拠バンドル: 2025 年第 1 四半期のパフォーマンス レポート、EMA で平滑化された指標の傾向 (6 か月のウィンドウ)、コスト関数分析、ベンダーのパフォーマンス レビュー、人口動態の影響分析。 4. 根拠: 「SFHAP は、3 四半期連続で 4 つの指標のうち 3 つの指標でパフォーマンスを下回りました。掲載枠あたりのコストは上昇傾向にあり、安定化の兆しはありません。ベンダー契約の再交渉、資格基準の改訂、およびプログラムの再設計の可能性を評価するために一時停止します。予想される一時停止期間: 90 日。 5. 通知: 受益者通知レターは 2025 年 2 月 20 日に郵送されました。 2025 年 2 月 21 日にスプリングフィールド官報および市ウェブサイトに公告が掲載されました。Council briefed at regular session 2025-02-22. 6. Checkpoint: CP-2025-0301-HOU-SFHAP created 2025-03-01. Integrity hash: SHA-256(S_B||S_F||S_O||S_M||...) = 0x7a3f...c812.
10.6 Pause Period Activities
During the 90-day pause (March-May 2025), the Housing Services department conducted the following evaluation activities:
- Vendor audit: Discovered that the primary housing placement vendor had subcontracted to a firm with a 42% placement failure rate, explaining the declining stability rate.
- 適格性分析: 収入基準値 (60% AMI) とスプリングフィールドの住宅市場を組み合わせると、適格な家族と利用可能なユニットの間に不一致が生じ、資本ギャップの一因となることが判明しました。
- Program redesign: Developed a revised program model with (a) a new vendor procurement, (b) adjusted eligibility to 50% AMI with a supplemental tier at 50-70% AMI, and (c) a housing stability support component (case management for the first 6 months post-placement).
10.7 再開の決定
2025年5月20日、評価委員会はパラメータを修正した再開を勧告した。履歴書の責任の連鎖:
1. Corrective action verification: New vendor contract signed (Blue River Housing, 91% historical stability rate). Eligibility criteria revised. Case management component designed and staffed. 2. Authority: City Manager Robert Torres, approved 2025-05-22. 3. Modified parameters: Vendor = Blue River Housing; eligibility = 50% AMI (primary) + 50-70% AMI (supplemental); case management = 6 months post-placement; revised budget = $4.5M/year (incremental $300K for case management). 4. Checkpoint restore: CP-2025-0301-HOU-SFHAP restored. 194 active beneficiaries re-engaged. Financial accounts reconciled.
10.8 再開後のパフォーマンス
The resumed SFHAP entered a 30-day intensive monitoring period (June 2025) followed by regular quarterly evaluation. Post-resumption performance:
|四半期 |家族が住んでいる |安定率 |コスト/配置 |資本ギャップ |
|---|---|---|---|---|
| 2025 年第 3 四半期 | 91 | 89% | $12,400 | 6% |
| Q4 2025 | 94 | 91% | $11,800 | 5% |
| Q1 2026 | 97 | 92% | $11,200 | 4% |
All four metric dimensions returned to within-target performance by Q4 2025, two quarters after resumption. The cost per placement decreased from $19,100 (pre-pause) to $11,200 (Q1 2026), a 41% improvement. The housing stability rate increased from 68% to 92%, a 24-percentage-point improvement. The equity gap closed from 18% to 4%, well within the 10% tolerance.
10.9 反事実分析
一時停止の枠組みがなかったら、どうなっていただろうか?一時停止前の軌跡と同様のプログラムの歴史的な前例に基づいて、次のようになります。
シナリオ A (従来のガバナンス): 2025 年 12 月の年次評価では、パフォーマンス不足が特定されていたでしょう。 2026年第1四半期の法的見直しでは、継続か終了かが議論されることになるだろう。受益者擁護派からの政治的圧力があれば、このプログラムは多少の変更を加えて継続されただろう。 2025 年 3 月から 2026 年 3 月までの追加の無駄の合計: 目標の 41% の成果をもたらすプログラムの運用支出として約 450 万ドル。
シナリオ B (サンセット条項): 3 年間のサンセット条項は 2027 年 1 月に発動されるはずです。プログラムは強制評価の前にさらに 22 か月間実行されるはずです。追加の無駄の合計: 約 770 万ドル。
Scenario C (Pausable Policy Design): The pause was triggered in February 2025, 6 months after performance deterioration began. The 90-day pause cost $524K. The resumed program achieved target performance within 2 quarters. Total cost of the intervention: $524K + $300K/year incremental budget = $824K. Net savings vs. Scenario A: $3.7M. Net savings vs. Scenario B: $6.9M.
11. ベンチマーク
We evaluate Pausable Policy Design against three baselines: traditional annual review governance, sunset clause governance, and the PPD framework implemented on MARIA OS. The evaluation uses a portfolio of 50 simulated municipal policies over a 5-year period, with varying performance trajectories (25% consistently performing, 35% gradually deteriorating, 25% fluctuating, 15% catastrophically failing).
11.1 Failing Policy Detection Rate
| Governance Model | Detection Rate | Mean Time to Detection | False Positive Rate |
|---|---|---|---|
|年次レビュー | 67.3% | 14.2ヶ月 | 2.1% |
|サンセット条項 (3 年) | 78.1% | 22.6ヶ月 | 0.8% |
| PPD (tau_pause = 0.5) | 94.7% | 4.8 months | 6.3% |
| PPD (tau_pause = 0.6) | 89.2% | 6.1ヶ月 | 3.1% |
| PPD (tau_pause = 0.4) | 97.1% | 3.2 months | 11.8% |
PPD at the default threshold (tau_pause = 0.5) detects 94.7% of failing policies with a mean time to detection of 4.8 months -- 9.4 months faster than annual review and 17.8 months faster than sunset clauses. The higher false positive rate (6.3% vs. 2.1%) reflects the sensitivity tradeoff: earlier detection comes with more false alarms. However, false positive pauses are low-cost events (the policy is paused briefly, evaluated, and resumed) compared to the high cost of undetected failures.
11.2 Cumulative Waste Reduction
|ガバナンスモデル |失敗したプログラムへの総支出額 |廃棄物率 |削減 vs. ガバナンスなし |
|---|---|---|---|
|ガバナンスなし | 4,720万ドル | 100% | -- |
| Annual Review | $38.1M | 80.7% | 19.3% |
| Sunset Clause | $33.4M | 70.8% | 29.2% |
| PPD (tau_pause = 0.5) | 2,970万ドル | 62.9% | 37.1% |
PPD により、失敗したプログラムによる累積無駄がガバナンスなしの場合と比較して 37.1% 削減され、年次レビューと比べて 17.8 パーセントポイント、サンセット条項と比べて 7.9 パーセントポイント改善されました。コスト削減は、継続か終了かの選択を迫られるのではなく、早期発見と一時停止 (オプション価値の維持) 機能によって促進されます。
11.3 説明責任の帰属
| Governance Model | Decisions with Complete Accountability | Decisions with Partial Accountability | Decisions with No Attribution |
|---|---|---|---|
|年次レビュー | 71.4% | 18.3% | 10.3% |
| Sunset Clause | 82.6% | 12.1% | 5.3% |
| PPD | 99.2% | 0.6% | 0.2% |
PPD achieves 99.2% complete accountability attribution -- every pause, resume, terminate, and override decision has a traceable authority, evidence bundle, justification, and notification record. The 0.8% gap represents emergency pauses where retroactive documentation was completed within the 48-hour window. Under annual review governance, 10.3% of program decisions have no attribution at all -- the program was continued or modified without any documented decision-maker taking responsibility.
11.4 Resumption Integrity
| Metric | Value |
|---|---|
|作成されたチェックポイント | 127 |
|チェックポイントが復元されました | 68 |
|整合性ハッシュ検証の合格率 | 100% |
|受益者状態復元精度 | 98.6% |
|財務調整の正確性 | 99.8% |
| Mean time to full resumption | 12.3 business days |
| Beneficiary disruption incidents | 3 (out of 2,847 beneficiary-pause events) |
The checkpoint mechanism achieves 98.6% beneficiary state restoration accuracy, with the 1.4% gap attributable to beneficiaries who relocated or experienced eligibility changes during the pause period that were not captured in the checkpoint. Financial reconciliation accuracy of 99.8% confirms that the checkpoint captures fiscal state with near-perfect fidelity. The 3 beneficiary disruption incidents (0.1% of beneficiary-pause events) involved delayed re-notification due to outdated contact information.
12. Future Directions
12.1 Predictive Pause Triggers
The current framework triggers pauses based on observed metric deterioration -- it is reactive. A natural extension is predictive pause triggers that anticipate performance failure before it materializes in the metrics. Machine learning models trained on the metric trajectories of historically failing programs could produce early warning signals 2-4 months before the reactive pause condition fires. The challenge is balancing sensitivity (catching failures early) against specificity (avoiding false alarms that erode trust in the framework).
A predictive model P_predict(m, t) would estimate the probability that P_pause will exceed tau_pause within the next T months, given current metric vector m at time t. When P_predict exceeds a configured confidence threshold, the system would issue a pre-pause advisory -- not a formal pause, but a heightened monitoring state that increases metric collection frequency and triggers a preliminary cost function analysis.
12.2 Cross-Policy Correlation Analysis
Municipal policies do not operate in isolation. A housing subsidy program's performance may be affected by changes in the transportation program (affecting beneficiaries' access to employment), the education program (affecting family decisions about where to live), or the economic development program (affecting the availability of affordable rental units). The current framework evaluates each policy independently.
今後の作業では、政策ポートフォリオ間の因果関係を特定する政策間相関モデルを開発する必要があります。ポリシー A のメトリクスが悪化すると、モデルはその悪化が内生的 (ポリシー A 自身の設計によって引き起こされる) か外生的 (ポリシー B、C、および D によって作成された環境の変化によって引き起こされる) であるかを評価します。内因性の悪化は政策 A の一時停止を正当化します。外因性の悪化は相互作用する政策の調整された見直しを正当化します。
12.3 市民フィードバックの統合
現在の指標フレームワークは、管理データ (登録者数、支出記録、成果評価) に依存しています。政策が目的とする人々の声を直接反映するものではありません。今後の作業では、構造化された市民フィードバックを 5 番目の指標次元として一時停止条件関数に統合する必要があります。
Citizen feedback would be collected through standardized surveys, public comment systems, and community meeting transcripts processed by NLP. The feedback would be scored on dimensions of satisfaction, accessibility, fairness, and responsiveness, and weighted into the composite pause condition function with a dedicated weight w_citizen. The technical challenge is ensuring that feedback collection is representative and resistant to gaming.
12.4 Inter-Municipal Benchmarking
Municipalities implementing PPD on MARIA OS could benefit from inter-municipal benchmarking: comparing their policies' performance against similar policies in comparable municipalities. A housing subsidy program in Springfield that costs $12,000 per successful placement might be performing well relative to its own targets but poorly relative to comparable programs in similar-sized cities that achieve $8,000 per placement.
自治体間のベンチマークには、標準化された指標の定義、プライバシーを保護するデータ共有プロトコル、および状況要因 (住宅市場の状況、人口構成、規制環境) の慎重な制御が必要です。連合学習技術を利用すれば、地方自治体が生のプログラム データを共有する必要なくベンチマークを実行できる可能性があります。
12.5 適応閾値校正
一時停止しきい値 tau_pause は現在、静的なガバナンス パラメーターとして設定されています。今後の作業では、自治体の過去の一時停止精度に基づいて tau_pause を調整する 適応しきい値キャリブレーション を開発する必要があります。現在のしきい値によって誤検知の一時停止 (不必要な中断) が多すぎる場合、システムは tau_pause を徐々に増やす必要があります。偽陰性 (失敗の見逃し) が多すぎる場合、システムは tau_pause を減らす必要があります。
適応型キャリブレーションはメタ学習の一種であり、ガバナンス フレームワークはそれ自体を制御することを学習します。重要な制約は、しきい値の調整が透過的で監査可能であり、人間によるオーバーライドの影響を受けなければならないということです。システムはそれ自体の感度を黙って下げることはできません。すべてのしきい値調整は、標準的な責任チェーンを通過します。
12.6 憲法上の統合
将来の最も深い方向性は、PPD を自治体の憲法および憲章の枠組みと統合することです。多くの市憲章には、プログラムの評価、予算の監督、公的説明責任に関する規定が含まれており、これらは PPD の枠組みにおける制約として形式化できる可能性があります。たとえば、「年間 100 万ドルを超えるすべてのプログラムは 18 か月ごとに独立した評価を受ける必要がある」という憲章要件は、最大一時停止間隔の制約としてエンコードできます。P_pause が 18 か月以内に評価されなかった場合、システムはメトリクスのパフォーマンスに関係なく、必須のレビューをトリガーします。
This constitutional integration would transform PPD from an administrative tool into a governance infrastructure layer that implements the city's fundamental governance commitments as executable constraints.
13. 結論
この文書では、政府の政策の中断を正式で責任のある、元に戻せる操作にするための数学的フレームワークである一時停止可能なポリシーの設計について説明しました。このフレームワークは、止められない政策、つまり一時停止し、評価し、決定するためのガバナンスメカニズムが存在しないためにリソースを消費し続け、次善の結果を生み出すプログラムという根本的な問題に対処します。
The key contributions are:
The Policy State Machine formalizes the policy lifecycle as a six-state automaton (Draft, Active, Paused, Resumed, Terminated, Completed) with guarded transitions that enforce governance requirements at the architectural level. The state machine provides the semantic foundation for treating policies as interruptible programs rather than irreversible commitments.
The Pause Condition Function P_pause(m) transforms observable performance metrics into a continuous urgency score, with weighted dimensions for effectiveness, efficiency, equity, and compliance. Temporal smoothing, hysteresis, and configurable thresholds prevent both over-sensitivity and under-sensitivity. The function makes the 'should we pause?' question answerable with quantitative evidence rather than political intuition.
The Accountability Requirement A(r) ensures that every pause decision has a traceable authority, an evidence bundle, a written justification, and stakeholder notification. The accountability chain is immutable, cryptographically hashed, and publicly accessible (subject to privacy protections). This addresses the accountability diffusion problem that is the root cause of policy persistence.
The Cost Function provides a three-way comparison (continue vs. pause vs. terminate) that makes the economic case for interruption explicit and auditable. The cost model includes operational expenditure, opportunity cost, harm cost, and political cost, enabling decision-makers to see the full picture rather than anchoring on sunk costs.
The Checkpoint Mechanism preserves policy state during pause, enabling resumption without beneficiary disruption or data loss. Checkpoints provide completeness, immutability, and restorability guarantees that make the pause reversible -- addressing the legitimate concern that pausing a program might destroy it.
民主的オーバーライド アーキテクチャ は、選出された役人の権限を維持しながら、枠組みの勧告を無効にする決定に対して強化された説明責任要件を課します。このフレームワークは意思決定支援システムであり、意思決定置換システムではありません。
このケーススタディは、PPD が MARIA OS に実装されていれば、SFHAP のパフォーマンス低下を従来の年次レビューよりも 10 か月早く検出し、累積無駄を 370 万ドル節約し、再開から 2 四半期以内に目標パフォーマンスを達成する再設計されたプログラムを生成できたであろうことを示しています。
ベンチマークでは、PPD が検出までの平均時間 4.8 か月 (年次レビューの場合は 14.2 か月) で失敗したポリシーの 94.7% を検出し、累積無駄を 37% 削減し、99.2% の説明責任の帰属を達成し、98.6% の再開の整合性を維持することが確認されています。
より広範な意味は、一時停止可能性はガバナンスの弱点ではなく、ガバナンスの能力であるということです。政策を一時停止できる政府は、間違いから学び、軌道修正し、より効果的に有権者に奉仕できる政府です。一時停止できない政府は存続するか破壊することしかできず、プログラムのパフォーマンスが低下しているものの回復可能な可能性がある場合、どちらの選択肢も公共の利益には役立ちません。
参考文献
- [1] 政府会計責任局。 (2024年)。 「2024 年年次報告書: 断片化、重複、重複を削減し、数十億ドルの経済的利益を達成する追加の機会。」 GAO-24-106915。プログラムの統合と終了による 5,210 億ドルの節約見積もりの主な情報源。
- [2] Pressman, J. および Wildavsky, A. (1984)。 「実装: ワシントンでの大きな期待がオークランドで打ち砕かれる方法」第3版カリフォルニア大学出版局。政府プログラムにおける政策設計と政策実行の間のギャップに関する古典的な分析。
- [3] Bardach, E. and Patashnik, E. (2019). "A Practical Guide for Policy Analysis: The Eightfold Path to More Effective Problem Solving." 6th ed. CQ Press. Standard framework for policy analysis including the evaluation criteria (effectiveness, efficiency, equity) that inform our metric dimensions.
- [4] Behn, R. (2014). "The PerformanceStat Potential: A Leadership Strategy for Producing Results." Brookings Institution Press. Analysis of performance management systems in government, including the challenges of metric-driven governance that motivate our temporal smoothing and hysteresis designs.
- [5] Moynihan, D. (2008). "The Dynamics of Performance Management: Constructing Information and Reform." Georgetown University Press. Research on how government organizations use (and fail to use) performance information, providing empirical grounding for the accountability gaming defenses.
- [6] Sunstein, C. (2014). "Simpler: The Future of Government." Simon & Schuster. Argument for evidence-based, adaptive government policies that aligns with the PPD framework's emphasis on continuous evaluation and formal pause conditions.
- [7] European Parliament. (2024). "Regulation (EU) 2024/1689 -- Artificial Intelligence Act." Official Journal of the European Union. Regulatory framework for AI governance that informs the transparency and accountability requirements of PPD.
- [8] National Institute of Standards and Technology. (2023). "AI Risk Management Framework (AI RMF 1.0)." NIST AI 100-1. US federal framework for AI governance, with accountability and transparency requirements that map to the PPD accountability chain.
- [9] Chandy, K.M. and Lamport, L. (1985). "Distributed Snapshots: Determining Global States of Distributed Systems." ACM Transactions on Computer Systems, 3(1), 63-75. The foundational algorithm for consistent snapshots in distributed systems, which inspires the checkpoint mechanism design.
- [10] Gray, J. and Reuter, A. (1993). "Transaction Processing: Concepts and Techniques." Morgan Kaufmann. Checkpoint and recovery theory from database systems, adapted for policy state preservation.
- [11] Argyris, C. および Schon, D. (1996)。 「組織学習 II: 理論、方法、実践」アディソン・ウェスリー。ポリシーの一時停止を組織の学習メカニズムとして扱うための概念的な基盤を提供する二重ループ学習理論。
- [12] フッド、C. (2011)。 「責任のゲーム: 政府におけるスピン、官僚主義、自己保存」プリンストン大学出版局。セクション 4.4 の説明責任ゲームの防御を動機付ける政府における責任回避行動の分析。
- [13] MARIA OS Technical Documentation. (2026). Internal architecture specification for the Decision Pipeline, Responsibility Gate Engine, Evidence Store, and MARIA Coordinate System.