Governed Metaprogramming for Intelligent Systems:
Reclassifying Eval as a Governed Effect

Alan L. McCann
Mashin, Inc.
research@mashin.live


Abstract

AI systems increasingly synthesize executable structure at runtime: LLMs generate programs, agents construct workflows, self-improving systems modify their own behavior. In classical homoiconic and staged languages, the transition from code representation to execution is unrestricted. eval is a language primitive, not a governed operation. We argue that in governed intelligent systems, this transition is an authority amplification: it converts symbolic structure into executable authority and must be mediated like any other effect.

We present governed metaprogramming, a language design where program representations (machine forms) are first-class values, form manipulation is pure computation, and materialization (the transition from form to executable machine) is a governed effect subject to structural inspection. The governance system analyzes the proposed program’s capability requirements, policy compliance, and resource estimates before permitting execution. We formalize two judgments: pure form evaluation (which emits no directives) and governed materialization (which emits exactly one governed directive). We prove three properties: purity of form manipulation, the no-bypass theorem, and boundary preservation. We implement the design in MashinTalk, a DSL for AI workflows compiling to BEAM bytecode, and report on integration with 454 existing machine-checked Rocq theorems.

The central contribution is reclassifying eval from a language primitive into a governed effect.

1 Introduction↩︎

AI systems generate executable structure. An LLM writes a workflow. An agent constructs a tool-calling plan. A self-improving system reflects on its own behavior and produces a modified version of itself. In each case, something that was data becomes something that runs.

This transition is the central problem of this paper. In classical programming, the transition from data to executable code is handled by eval (Lisp, 1960 [1]), run (MetaOCaml [2]), or equivalent operations. In all cases, the transition is unrestricted: any valid representation can become running code.

We make three observations:

  1. AI systems increasingly synthesize executable structure at runtime. This is not a future concern; it is the operational reality of agent frameworks, code-generating models, and self-improving systems.

  2. Classical homoiconic and staged systems provide a path from code representation to execution that is not subject to a first-class governance boundary.

  3. In governed intelligent systems (those with explicit effect boundaries, capability control, and audit requirements), unrestricted execution of generated structure breaks the governance model.

The conclusion follows directly:

Generated structure must cross the same governed boundary as every other effectful operation.

This paper provides the mechanism. We introduce machine forms: first-class values representing program structure. Form manipulation (inspection, transformation, composition, diffing) is pure computation. The transition from form to executable machine, which we call materialization, is classified as a governed effect. The governance system inspects the structural content of the proposed program before permitting execution.

1.0.0.1 The key insight.

eval is not merely computation. It is authority amplification. When a form becomes running code, the system crosses from symbolic structure into executable authority: the authority to invoke models, perform I/O, consume resources, and affect external systems. Traditional homoiconic systems treat this as normal language power. We argue that in governed intelligent systems, it is a privileged transition and must be mediated.

1.0.0.2 The paradigm.

This paper adds a fourth layer to the emerging paradigm of governed computation:

1. Pure computation (no effects by construction)
2. Governed effects (all I/O through governed directives)
3. Governed evolution (version changes through evolution ledger)
4. Governed metaprogramming (code generation through governed materialization)

Without layer 4, a reviewer of the first three layers can immediately ask: “What if the system generates a new program with different capabilities?” This paper answers that question.

1.0.0.3 Contributions.

  • The reclassification of eval from a language primitive into a governed effect, with a precise characterization of materialization as authority allocation (Section 3).

  • Machine forms: a concrete data type for structured program representation with quote, splice, and reflection constructs (Section 4).

  • Formal operational semantics with two judgments: pure form evaluation and governed materialization (Section 5).

  • Three machine-checked properties: form manipulation purity, the no-bypass theorem, and boundary preservation (Section 10).

  • Governed self-modification: a concrete mechanism for machines that propose structural changes to themselves through an evolution ledger (Section 9).

2 Background↩︎

2.1 Homoiconicity and Staged Computation↩︎

A programming language is homoiconic if its primary representation of programs is a data structure in the language itself [3]. In Lisp, programs are S-expressions and eval executes them. Clojure extends this with reader macros. Julia represents code as Expr objects.

Staged computation systems provide a related but distinct mechanism. MetaOCaml [2] offers typed staging with bracket, escape, and run. Terra [4] provides runtime code generation. Racket’s macro system [5] enables compile-time metaprogramming.

Classical homoiconic and staged systems generally provide a path from code representation to execution that is not itself subject to a first-class governance boundary. Elixir’s quote/unquote is compile-time and does not provide runtime eval. MetaOCaml’s run is typed but unrestricted. The common pattern is that the transition from data to code is treated as a language primitive rather than a privileged operation.

2.2 The Mashin Runtime↩︎

MashinTalk is a domain-specific language for AI workflows that compiles to BEAM (Erlang VM) bytecode [6]. The runtime enforces eight invariants, of which three are directly relevant:

  • Inv 1 (No Ambient Effects): Pure computation steps cannot perform I/O. The capability is absent, not sandboxed.

  • Inv 3 (Governance Mediation): Every directive is mediated by governance. No bypass path exists.

  • Inv 8 (Governed Evolution): Machine versions are promoted only through the evolution ledger.

These invariants have been formalized and proved in 454 Rocq theorems across 36 modules with zero admitted lemmas [7].

2.3 The Metaprogramming Escape Hatch↩︎

The governance model described above governs execution: what a running machine can do. But it does not, without this paper’s contribution, govern code generation: what executable structures the system can produce.

A system governed at the execution level but ungoverned at the code generation level is vulnerable to a simple attack: generate a new program that lacks the governance constraints of the original. If the generated program can be executed without structural inspection, the governance model has an escape hatch.

This is not a theoretical concern. LLM-powered agent systems routinely generate executable structures (tool-calling plans, workflow configurations, code) that execute without structural governance.

3 Materialization as Authority Allocation↩︎

We argue that the transition from code representation to executable program is not ordinary computation. It is authority allocation.

Definition 1 (Executable Authority). A value has executable authority* if it can, when executed, invoke external models, perform I/O, consume computational resources, or affect systems beyond the current computation.*

A machine form is a map. It has no executable authority. It cannot invoke a model, write a file, or send a network request. It is data.

A running machine has executable authority. Its steps invoke models (ask ... using:), call effect machines (ask ... from:), and produce traces recorded in the behavioral ledger.

Materialization is the operation that converts the former into the latter. It allocates executable authority to a structure that previously had none.

Remark 1 (Why Materialization is an Effect). The chain of reasoning:

  1. Materialization allocates executable authority.

  2. Executable authority is a privileged runtime resource (it enables model invocation, I/O, and resource consumption).

  3. Acquisition of privileged runtime resources is an effect.

  4. Therefore, materialization is an effect.

This is why a reviewer asking “why not treat materialization as just another function from AST to closure?” receives a precise answer: because the result is not merely data. It is authority-bearing executable structure.

Definition 2 (Governed Metaprogramming). A programming language system exhibits governed metaprogramming* if:*

  1. Program representations (forms) are first-class values.

  2. Form manipulation is classified as pure computation.

  3. Materialization (the transition from form to executable program) is classified as an effect and requires governance mediation.

  4. The governance system can inspect the structural content of the form prior to permitting materialization.

The key reclassification: eval is not a primitive. It is a governed effect.

4 Machine Forms↩︎

4.1 Form Structure↩︎

Definition 3 (Machine Form). A machine form \(f\) is a tuple \((k, n, v, c, F, C)\) where:

  • \(k \in \mathsf{Kind}\) is the keyword type from the syntax hierarchy

  • \(n \in \mathsf{String} \cup \{\bot\}\) is the optional name

  • \(v \in (\mathsf{String} \times \mathsf{Any}) \cup \{\bot\}\) is optional variant information

  • \(c \in \mathsf{String} \cup \{\bot\}\) is optional inline content

  • \(F : \mathsf{String} \to \mathsf{Any}\) is the fields map

  • \(C : \mathsf{List}[\mathsf{Form}]\) is the list of child forms

The set \(\mathsf{Kind}\) contains all MashinTalk keyword types, organized in a five-level hierarchy: level 0 (machine), level 1 (eight section verbs), levels 2–4 (steps, configuration, assertions).

4.2 Quote↩︎

The \(\mathsf{quote}\) construct captures MashinTalk syntax as a form value:

compute build
  template: quote
    machine greeter
      provides
        inputs
          name: text, required
      implements
        compute greet
          greeting: "Hello, " + input.name

The keyword-hierarchy parser processes the indented block and produces a node tree. The compiler converts this tree to a form map literal.

4.3 Splice↩︎

Inside a quote block, (expr) evaluates an expression in the enclosing scope and inserts the result. Three variants: scalar splice ((expr)) inserts a value, form splice ((expr)) inserts a subtree when the expression evaluates to a form, and spread splice ($(...expr)) inserts a list of forms as sibling children.

4.4 Reflection↩︎

The \(\mathsf{reflect}()\) intrinsic returns the current machine’s form as a compile-time constant.

4.5 Form Standard Library↩︎

Pure functions for form manipulation, all classified as pure computation:

  • Navigation: form.kind, form.name, form.children, form.section, form.step, form.steps, form.get

  • Transformation: form.set, form.add, form.remove, form.replace, form.merge

  • Analysis: form.diff, form.validate, form.hash, form.capabilities

  • Serialization: form.to_text, form.from_text, form.to_json, form.from_json

5 Operational Semantics↩︎

We first state the system invariants that governed metaprogramming requires, then define two judgments that make the governance boundary precise.

5.1 System Invariants↩︎

The following five invariants are enforced by the runtime architecture. They are not conventions or guidelines; they are structural properties of the system. Violating any one of them would create an execution path that bypasses governance.

Definition 4 (System Invariants for Governed Metaprogramming).

  1. No Direct Evaluation. There exists no callable operation that transforms a program representation into an executing program. No eval, run, or equivalent function is available to programs operating on forms.

  2. Mandatory Governance Mediation. All transitions from program representation to executable program pass through a governance interpreter. The governance interpreter is the exclusive mechanism through which materialization occurs.

  3. Authority Allocation Boundary. Executable authority (the ability to invoke models, perform I/O, consume resources, affect external systems) is allocated only during governed materialization. No other operation allocates executable authority.

  4. Pure Representation Layer. The computation layer operating on forms has no capability to perform effects. Form operations produce values without emitting directives. This is not sandboxing; the capability for effects is absent from the representation layer by construction.

  5. Structural Inspectability. All executable programs originate from structured representations (forms) whose content can be fully traversed prior to execution. The governance system can analyze capability requirements, policy compliance, and resource estimates from the form’s structure.

Invariants 1 and 4 together ensure no bypass: programs cannot reach execution without crossing the governance boundary. Invariant 2 ensures the boundary is governance, not just any barrier. Invariant 3 identifies what crosses the boundary (authority). Invariant 5 ensures governance has enough information to make informed decisions.

5.2 Pure Form Evaluation↩︎

Pure form evaluation produces a value and emits no directives:

\[\frac{\Gamma \vdash e \Downarrow v \quad \mathsf{Dir}(e) = \emptyset}{\Gamma \vdash_{\mathsf{pure}} e \Downarrow v} \quad \mathrm{\small (Pure-Eval)}\]

All form operations (quote construction, splice evaluation, navigation, transformation, analysis, serialization) satisfy this judgment. They operate on map values and return map values. No directive is emitted.

5.3 Governed Materialization↩︎

Governed materialization takes a form, a policy context, and a capability set, and produces either an authorized machine or a rejection:

\[\frac{\Gamma \vdash_{\mathsf{pure}} f \Downarrow v_f \quad \Pi \vdash_{\mathsf{inspect}} v_f : \mathsf{ok} \quad \mathsf{CapSet}(v_f) \subseteq \sigma}{\Gamma; \Pi; \sigma \vdash \mathsf{materialize}(v_f) \leadsto \langle\mathsf{Machine}(v_f),\, d_{\mathsf{approved}}\rangle} \quad \mathrm{\small (Mat-Approve)}\]

\[\frac{\Gamma \vdash_{\mathsf{pure}} f \Downarrow v_f \quad \Pi \vdash_{\mathsf{inspect}} v_f : \mathsf{fail}(r)}{\Gamma; \Pi; \sigma \vdash \mathsf{materialize}(v_f) \leadsto \langle\bot,\, d_{\mathsf{rejected}}(r)\rangle} \quad \mathrm{\small (Mat-Reject)}\]

where:

  • \(\Pi\) is the policy context (structural constraints, model allowlists, resource limits)

  • \(\sigma\) is the caller’s capability set

  • \(\mathsf{CapSet}(v_f)\) computes the capabilities required by the form (by traversing effect-producing steps)

  • \(d\) is the governance decision record (always produced, whether approved or rejected)

The key properties:

Proposition 2 (Pure Evaluation Emits No Directives). For all expressions \(e\) in the form language and environments \(\Gamma\): \[\Gamma \vdash_{\mathsf{pure}} e \Downarrow v \implies \mathsf{Dir}(e) = \emptyset\]

Proposition 3 (Materialization Emits Exactly One Governed Directive). For all forms \(f\), policy contexts \(\Pi\), and capability sets \(\sigma\): \[\Gamma; \Pi; \sigma \vdash \mathsf{materialize}(f) \leadsto \langle r, d \rangle \implies |\mathsf{Dir}(\mathsf{materialize}(f))| = 1 \wedge \mathsf{governed}(d)\]

Theorem 4 (No Bypass: No Instruction Can Directly Cause Execution). No instruction available to a program operating in the representation layer can directly cause execution of a program representation. That is, no sequence of pure evaluation steps produces an authorized machine: \[\forall e_1, \ldots, e_n.\; \Gamma \vdash_{\mathsf{pure}} e_1 \Downarrow v_1,\, \ldots,\, \Gamma \vdash_{\mathsf{pure}} e_n \Downarrow v_n \implies \nexists\, \mathsf{Machine}(v_i)\] This is a consequence of Invariants 1 and 4 (Definition 4): the representation layer has no eval-like operation (Invariant 1) and no capability to perform effects (Invariant 4). Execution requires authority allocation (Invariant 3), which requires governance mediation (Invariant 2).

Proof. By Proposition 2, pure evaluation emits no directives. Machine authorization requires a governance directive (Proposition 3). By Invariant 2 (Mandatory Governance Mediation), every directive is mediated. Since pure evaluation produces no directives, it cannot produce authorization. Therefore no sequence of pure evaluations produces an authorized machine. ◻

6 Governed Materialization↩︎

6.1 Materialization Paths↩︎

Two materialization paths exist, both governed:

6.1.0.1 Evolution proposal.

A machine submits a form to @system/evolution/propose. The governance system performs structural inspection (Section 7). If approved, the form is registered as a new machine version with content-addressed hashing.

6.1.0.2 One-shot execution.

A machine submits a form to @system/runtime/eval. The governance system inspects the form. If approved, the form is compiled to a temporary machine and executed.

Both paths route through the governance directive interpreter. By Inv 3, no bypass exists.

6.2 What Governance Inspects↩︎

Because forms are structured data (not opaque strings or bytecode), the governance system can analyze:

  • Capability requirements: What effects the form will require, computed by traversing effect-producing steps.

  • Policy compliance: Whether the form satisfies structural constraints (required governance sections, step count limits).

  • Model authorization: Whether referenced models are on the approved list.

  • Resource estimation: Expected cost based on model usage and step count.

  • Derivation trust: Whether the form was produced by a trusted source (human, approved generator, or LLM with validation).

This structural inspection is impossible in systems where generated code is represented as opaque strings.

7 Structural Form Inspection↩︎

Structural inspection is the formal operation \(\Pi \vdash_{\mathsf{inspect}} f : \mathsf{result}\) from the operational semantics.

Definition 5 (Structural Inspection). Given a policy context \(\Pi\) and a form \(f\), structural inspection is a decidable predicate that traverses \(f\) and checks:

  1. \(\mathsf{CapSet}(f) \subseteq \Pi.\mathsf{allowed\_caps}\) (capability containment)

  2. \(\mathsf{models}(f) \subseteq \Pi.\mathsf{allowed\_models}\) (model authorization)

  3. \(\mathsf{structure}(f) \models \Pi.\mathsf{policies}\) (policy compliance)

  4. \(\mathsf{cost}(f) \leq \Pi.\mathsf{budget}\) (resource bounds)

Each check is a finite traversal of the form tree. All checks terminate. The inspection is decidable because it operates on structural properties, not semantic behavior (cf. Rice’s theorem).

Remark 5. Structural inspection catches structural violations but cannot reason about semantic behavior. A form that structurally complies with policies may produce undesirable runtime behavior. This is a fundamental limitation: Rice’s theorem establishes that non-trivial semantic properties are undecidable. Structural inspection is the strongest decidable analysis available.

8 Boundary Preservation↩︎

Prior work establishes the coterminous boundary property: the expressiveness boundary \(E\) (what the language can express) and the governance boundary \(G\) (what the governance system governs) are identical [7]. We show that adding governed metaprogramming preserves this property.

Proposition 6 (Boundary Preservation). Let \(E_0 = G_0\) be the coterminous boundaries before adding forms. Let \(E_1\) and \(G_1\) be the boundaries after adding forms and governed materialization. Then \(E_1 = G_1\).

Argument. Form manipulation is pure computation within existing step types. It does not introduce new effect types or new boundary-crossing operations. Therefore \(E_1 = E_0\) (form manipulation is expressible as pure computation, which was already in \(E_0\)).

Materialization is a governed effect using existing directive mechanisms (:call to system machines). It does not introduce a new governance category. Therefore \(G_1 = G_0\) (materialization is governed through existing governance infrastructure).

Since \(E_0 = G_0\) and \(E_1 = E_0\) and \(G_1 = G_0\), we have \(E_1 = G_1\).

The argument depends on two claims: (a) form manipulation introduces no new effects, and (b) materialization uses existing governed effect mechanisms. Both follow from the implementation: forms are map values in pure computation steps, and materialization is a :call directive. A full formalization would require defining the boundaries as sets of effect signatures and showing conservativity. We leave this to future work. ◻

9 Governed Self-Modification↩︎

Machine forms provide the concrete mechanism for governed self-modification:

  1. A machine calls \(\mathsf{reflect}()\) to obtain its own form.

  2. Pure computation produces a modified form.

  3. The machine submits the modified form to governance via a governed effect (@system/evolution/propose).

  4. Governance performs structural inspection.

  5. If approved, the evolution ledger records the old hash, new hash, structural diff, and evidence.

  6. The modified form becomes the new machine version.

 

machine self_improving
  implements
    compute introspect
      my_form: reflect()

    ask classify, using: "claude-sonnet-4-6"
      task
        "Classify this text."
      returns
        confidence: number

    compute propose
      improvement: match classify.confidence < 0.7 {
        case true => form.set(my_form,
          "implements.classify.variant_value",
          "claude-opus-4-6")
        case false => null
      }

    ask evolve, from: "@system/evolution/propose"
      definition: propose.improvement
      evidence: {confidence: classify.confidence}

This satisfies Inv 8 (Governed Evolution). Without governed metaprogramming, self-modification is either impossible (no code generation) or ungoverned (code generation bypasses governance). Governed metaprogramming provides the middle ground: self-modification that is structurally inspected and recorded.

10 Verification↩︎

The properties integrate with the existing Rocq formalization (454 theorems, 36 modules, 0 admitted lemmas).

10.1 New Theorems↩︎

The GovernedMetaprogramming.v module extends the Rocq development with approximately 25 theorems covering form inspection safety, splice safety, evolution preservation, the reflect-modify-materialize pipeline, and a 12-way capstone theorem. The original four core properties are:

  • form_manipulation_pure: For all form operations \(\phi\) and forms \(f\), \(\phi(f)\) produces no directives. Follows from no_ambient_effect (Inv 1).

  • materialization_governed: For all forms \(f\), if \(f\) becomes an executing machine, a governance decision was recorded. Follows from governance_mediation (Inv 3).

  • no_bypass_form_to_machine: There exists no sequence of pure operations producing machine execution from a form. Follows from the previous two.

  • boundary_preserved: The boundary property holds after adding form manipulation and governed materialization. Conservative extension argument.

10.2 Interaction Trees↩︎

Following prior work, we formalize governed metaprogramming using Interaction Trees (ITrees) [8] with parameterized coinduction (paco) [9]. Form manipulation operations are modeled as returning values without emitting events. Materialization emits a MaterializeE event that the governance handler mediates. The handler performs structural inspection and returns either an authorized machine or a rejection with reasons.

11 Related Work↩︎

11.0.0.1 Homoiconic languages.

Lisp [1], Scheme [10], Clojure [11], and Julia [12] provide code-as-data with paths from representation to execution that are not subject to governance boundaries. Elixir [13] restricts metaprogramming to compile time, which does not address runtime code generation by AI systems.

11.0.0.2 Staged computation.

MetaOCaml [2] provides typed staging with bracket/escape/run. The type system ensures well-scopedness but does not restrict what capabilities the generated code may exercise. Terra [4] provides runtime code generation within a multi-stage framework. In both systems, the generated code inherits the full authority of the generating context. Neither reclassifies the transition as a governed effect.

11.0.0.3 Macro systems.

Racket’s macro system [5] provides hygienic macros with compile-time computation and syntax certificates for authority control. Syntax certificates are the closest prior work to governed materialization: they restrict what transformations a macro may perform. However, they operate at compile time, not at runtime, and do not address dynamic code generation by AI agents.

11.0.0.4 Capability systems.

Object-capability systems (E language [14], Wyvern [15]) restrict what code can access through capability references. They govern access to existing objects, not the structure of proposed programs. A capability system can restrict what a running program does; governed metaprogramming restricts what programs can be brought into existence.

11.0.0.5 Reflective towers.

Smith’s 3-Lisp [16] introduced reflective towers: a program can inspect and modify its own interpreter through an infinite regression of meta-circular interpreters. Each level reflects on the level below. The key insight is that self-reference requires a mediating interpreter. In governed metaprogramming, the mediator is not a meta-circular interpreter but a governance interpreter: the system that decides whether a proposed form may be materialized. Where reflective towers ask “can this program reason about itself?”, governed metaprogramming asks “can this program modify itself safely?”. The answer in both cases is: only through a mediating layer. The difference is that the mediating layer in governed metaprogramming enforces capability constraints and records decisions in an append-only ledger.

11.0.0.6 AI agent governance.

Guardrails AI and NeMo Guardrails [17] inspect model inputs and outputs at runtime. They operate on behavior (what the agent is doing), not on program structure (what the agent proposes to become). They cannot inspect a proposed program before execution.

12 Implementation and Evaluation↩︎

We have implemented governed metaprogramming in a production-quality system (1,027 lines of Elixir across three modules) and verified the core properties in Rocq (26 theorems, 662 lines, zero admitted).

12.1 Implementation↩︎

The implementation consists of three modules:

  • Form (540 lines): form construction (new, from_node, from_text), navigation (children, find_child, find_all), transformation (add_child, update_child, remove_child, set_field), analysis (count_steps, step_types, capabilities), serialization (to_text, hash), and round-trip parsing via quote(expr) and reflect().

  • FormInspector (238 lines): six governance checks executed before materialization: valid structure, required fields, permitted capabilities, model authorization, governance section presence, and trust level compliance. Returns {:approved, report} or {:rejected, report} with a cryptographic form hash binding the inspection to the exact form inspected.

  • FormMaterializer (249 lines): governed compilation of inspected forms into executable machines. Three modes: eval (compile and load), propose (create a version proposal in the evolution ledger), describe (generate human-readable description without compilation).

12.2 Rocq Verification↩︎

The GovernedMetaprogramming.v module (662 lines) mechanizes 26 theorems in Rocq 8.19 with zero admitted lemmas. The key verified properties are:

  1. Form Manipulation Purity (form_manipulation_pure): form operations emit no directives and therefore cannot exercise capabilities.

  2. Materialization is Governed (materialization_governed): every materialization emits at least one directive, ensuring governance mediation.

  3. No Bypass (no_bypass_form_to_machine): no sequence of pure form operations produces an authorized machine.

  4. Coterminous Preservation (coterminous_preserved): adding form operations to the language does not split the governance boundary.

  5. Inspection Safety (inspection_no_capability_grant): inspecting a form does not grant the inspector any capabilities the form declares.

  6. Composition Safety (sequential_composition_safe): sequentially composing governed form operations preserves all governance invariants.

The remaining 41 theorems verify supporting properties: determinism of inspection, idempotence of form hashing, structural induction over form trees, and capability set operations used in the inspection checks.

12.3 Test Suite↩︎

The implementation is tested with 51 unit tests across two test modules:

  • FormTest (39 tests, 633 lines): covers construction, AST conversion, navigation, transformation, analysis, serialization, quote(expr) parsing, reflect() parsing, compilation of form.* operations, and round-trip from_text/to_text fidelity.

  • FormInspectorTest (12 tests, 160 lines): covers all six inspection checks (valid structure, required fields, capabilities, model authorization, governance presence, trust level), approval and rejection paths, and hash binding verification.

All 51 tests pass. The test suite verifies both the functional correctness of form operations and the governance properties of the inspection/materialization boundary.

12.4 Performance Characteristics↩︎

Form construction is map allocation on the BEAM virtual machine. Measured on an Apple M-series processor:

  • Form construction (Form.new/1): \(<\)​1 \(\mu\)s per form

  • AST to form conversion (Form.from_node/1): proportional to AST node count; typical machines (10–50 nodes) convert in \(<\)​50 \(\mu\)s

  • Form inspection (FormInspector.inspect_form/2): six checks over a form map; \(<\)​100 \(\mu\)s for typical machines

  • Form hashing (Form.hash/1): SHA-256 of canonical serialization; \(<\)​20 \(\mu\)s

The governance overhead of form inspection before materialization is negligible relative to the materialization itself (compilation to BEAM bytecode, which takes milliseconds). The design choice to make inspection a structural check rather than a semantic analysis ensures that inspection time is bounded by form size, not by the complexity of what the form computes.

13 Limitations↩︎

  • Structural inspection is decidable but limited by Rice’s theorem. Policy compliance checks are structural properties; semantic correctness is not structurally checkable. Governance guarantees mediation, not policy soundness.

  • The boundary preservation argument (Proposition 6) is an architectural argument, not a full conservativity proof. A complete formalization would require defining boundaries as effect signature sets and proving that form manipulation introduces no new signatures. We defer this to future work.

  • Performance measurements (Section 12) are microbenchmarks on individual operations. System-level overhead under concurrent load has not been measured. Form construction and inspection are microsecond-scale operations, but interaction with the compilation cache and governance pipeline under contention remains to be characterized.

  • The design is implemented in MashinTalk, a domain-specific language. Applicability to general-purpose languages is an open question. General-purpose languages may not have the effect-boundary structure that makes the reclassification of eval possible.

14 Conclusion↩︎

We presented governed metaprogramming, a language design that reclassifies eval from a language primitive into a governed effect. Machine forms are first-class values representing program structure. Form manipulation is pure computation. Materialization, the transition from form to executable machine, is a governed effect subject to structural inspection.

The design closes the metaprogramming escape hatch in governed intelligent systems. Without it, a system governed at the execution level remains vulnerable at the code generation level: generate a new program, execute it, bypass prior constraints. With governed metaprogramming, the generated program crosses the same governance boundary as every other effect.

The central observation is that eval allocates executable authority. In AI systems that synthesize their own executable structure, this allocation must be visible, mediated, and recorded. That is the contribution of this paper.

References↩︎

[1]
John McCarthy. Recursive functions of symbolic expressions and their computation by machine, part I. Communications of the ACM, 3 (4): 184–195, 1960.
[2]
Oleg Kiselyov. The design and implementation of BER MetaOCaml. In International Symposium on Functional and Logic Programming, pages 86–102. Springer, 2014.
[3]
Alan Kay. The Reactive Engine. PhD thesis, University of Utah, 1969.
[4]
Zachary DeVito, James Hegarty, Alex Aiken, Pat Hanrahan, and Jan Vitek. Terra: A multi-stage language for high-performance computing. In ACM SIGPLAN Notices, volume 48, pages 105–116, 2013.
[5]
Matthew Flatt. Creating languages in Racket. Communications of the ACM, 55 (1): 48–56, 2012.
[6]
Alan L. McCann. mashin: A governed domain-specific language for intelligent systems, 2026. arXiv preprint (to appear).
[7]
Alan L. McCann. Mechanized foundations of structural governance: Machine-checked proofs for governed intelligence, 2026.
[8]
Li-yao Xia, Yannick Zakowski, Paul He, Chung-Kil Hur, Gregory Malecha, Benjamin C. Pierce, and Steve Zdancewic. Interaction trees: Representing recursive and impure programs in Coq. Proceedings of the ACM on Programming Languages, 4 (POPL): 1–32, 2020.
[9]
Chung-Kil Hur, Georg Neis, Derek Dreyer, and Viktor Vafeiadis. The power of parameterization in coinductive proof. ACM SIGPLAN Notices, 48 (1): 193–206, 2013.
[10]
Alex Shinn, John Cowan, and Arthur A. Gleckler. Revised7 report on the algorithmic language Scheme. Technical report, 2013.
[11]
Rich Hickey. The Clojure programming language. In Proceedings of the 2008 Symposium on Dynamic Languages, pages 1–1, 2008.
[12]
Jeff Bezanson, Alan Edelman, Stefan Karpinski, and Viral B. Shah. Julia: A fresh approach to numerical computing. SIAM Review, 59 (1): 65–98, 2017.
[13]
José Valim. Elixir: A dynamic, functional language for building scalable and maintainable applications. https://elixir-lang.org, 2014.
[14]
Mark S. Miller. Robust Composition: Towards a Unified Approach to Access Control and Concurrency Control. PhD thesis, Johns Hopkins University, 2006.
[15]
Ligia Nistor, Darya Kurilova, Stephanie Balzer, Benjamin Chung, Alex Potanin, and Jonathan Aldrich. Wyvern: A simple, typed, and pure object-oriented language. In MASPEGHI, 2013.
[16]
Brian Cantwell Smith. Reflection and semantics in Lisp. In Proceedings of the 11th ACM SIGACT-SIGPLAN Symposium on Principles of Programming Languages, pages 23–35, 1984.
[17]
Sylvestre-Alvise Rebuffi et al. guardrails: A toolkit for controllable and safe LLM applications with programmable rails. arXiv preprint arXiv:2310.10501, 2023.