"""
reify — FABRIC inverse pass: JSON inference → regenerated text

The vivify pipeline moves from language to structure: raw text enters, keywords
and clumps emerge, the inference is filed into a category tree, and tension
is scored over the operator coordinates. reify runs that process in reverse. It
takes a stored inference — already compressed into left_keywords, clumps,
category_paths, and its tension block — and asks the Claude API to reconstruct the
analog original: the felt thought the structure was built from.

This is not summarization or paraphrase. The model is instructed to speak from
inside the meaning, not about it.

Tension shapes the output, but NOT in the way this file claimed until 2026-09-19.
The old prompt described tension_score as "how far left and right keyword sets
diverge" and told the model to lean into that gap. That was the lexical score,
dead since 2026-07-13, and the number now carries a different meaning entirely:

  predicted   the operators' judgment of UN-TRUTH — the resonance surface vs
              underlying gap (illusion) blended with conflict alarms. High means
              the text presented one thing while something else sat underneath.
  confirmed   the same un-truth measured against the record, from claimed-vs-
              actual discrepancies. Only present on calibration material.
  None        unmeasured. NOT low, NOT sincere. Absence of a measurement.

So a high-tension reconstruction should keep the doubleness the original had —
surface intact, the underneath showing through — rather than lean into a
keyword-overlap gap that no longer exists. calibration_delta is deliberately NOT
sent: it measures operator error, not anything about the felt thought.

Three modes:

  single
    Reconstructs prose from one inference. The model receives left_keywords, clumps,
    and a slice of category_paths, plus the tension score. Output is 3-6 dense
    sentences in first person. Use this to test whether vivify actually captured
    what was meant — if the reconstruction feels foreign, the keywords drifted.

  synthesize
    Takes two inference files and generates a single passage that holds both
    simultaneously. Not alternating, not summarizing — finding the place where they
    are the same thought. If the two inferences pull in different directions the
    model writes from that tension. Use this to discover connections the corpus
    has not yet made explicit.

  voice
    Walks an entire category directory and generates a passage that speaks for the
    whole category — a distillation of what all its inferences are reaching toward
    together. This is the most generative mode: the emergent category tree, built
    by co-occurrence across the full corpus, becomes the source material for new
    prose that could not have come from any single inference.

Options:
  --dry-run    Print the full prompt without calling the API. No billing. Use this
               to inspect what the model will receive before committing to a call.
  --dir        Inferences directory (default: inferences). Used with --voice.

Usage:
  python3 reify.py <path/to/inf_XXX.json>
  python3 reify.py --synthesize <path/to/inf_A.json> <path/to/inf_B.json>
  python3 reify.py --voice <category/path>
  python3 reify.py --dry-run <path/to/inf_XXX.json>
  python3 reify.py --dir <inferences_dir> --voice <category/path>
"""

import sys
import json
import argparse
from pathlib import Path

sys.path.insert(0, str(Path(__file__).parent / "lib"))
from vivify_core import read_json, resolve_model, llm_call
from category_index import ids_for_path, load_by_id


REIFY_PROMPT = """You are the inverse pass of a vivify pipeline.

You will receive a structured inference: left_keywords (felt semantic meaning),
clumps (grouped keyword clusters), category_paths (emergent filing), and a
tension block.

Your task is to reconstruct the felt thought — the analog original — from this
structure. Do not explain the keywords. Do not describe what the pipeline did.
Speak from inside the meaning, not about it.

Rules:
- Write in first person, direct voice
- Do not mention keywords by name — let them shape the prose, not appear in it
- Do not reference the pipeline, categories, or JSON
- The tension block measures UN-TRUTH, not keyword overlap:
  - "predicted" (0-1) is how far the text's surface stood apart from what was
    underneath it. High means the original said one thing and meant, or was, another.
    Reconstruct that doubleness — let the surface stand while the underneath shows
    through it. Do not resolve the gap and do not announce it.
  - "confirmed" (0-1) is the same distance measured against the record rather than
    inferred. When present and high, the gap is established fact, not suspicion;
    write with that footing.
  - null, or a missing tension block, means UNMEASURED. It does not mean low, and
    it does not mean the thought was sincere. Write without leaning either way.
- Output format: a series of sentence+bullets constructs. Each block is one
  direct claim sentence (active voice, 25 words max), followed by 3-5 bullet
  points that each expand a distinct angle — evidence, example, or constraint.
  No concluding statements. No bullets that restate the opener.
- Produce 3-5 such blocks. Each block covers a distinct facet of the felt meaning.

Inference:
"""

SYNTHESIZE_PROMPT = """You are the synthesis pass of a vivify pipeline.

You will receive two structured inferences. Each has left_keywords capturing its
felt meaning and clumps grouping those keywords. Your task is to generate output
that holds both inferences simultaneously — not alternating between them, not
summarizing them, but finding the place where they are the same thought.

Rules:
- Write in first person, direct voice
- Do not mention keywords by name
- Do not reference the pipeline, categories, or JSON
- If the two inferences pull in different directions, find the tension and write from it
- Output format: a series of sentence+bullets constructs. Each block is one direct
  claim sentence (active voice, 25 words max) followed by 3-5 bullets expanding
  distinct angles. Produce 3-5 blocks. The synthesis should feel inevitable, not constructed.

Inference A:
{inf_a}

Inference B:
{inf_b}
"""

VOICE_PROMPT = """You are the category voice pass of a vivify pipeline.

You will receive a set of inferences that share a category. Each has left_keywords
capturing its felt meaning. Your task is to speak for the whole category — a
distillation of what all these inferences are reaching toward together.

Rules:
- Write in first person, direct voice
- Do not mention keywords by name
- Do not reference the pipeline, categories, or JSON
- This is not a summary — it is a synthesis. Find the irreducible core.
- Output format: a series of sentence+bullets constructs. Each block is one direct
  claim sentence (active voice, 25 words max) followed by 3-5 bullets expanding
  distinct angles. Produce 3-5 blocks.

Category: {category}

Inferences:
{inferences}
"""


def _tension_for_prompt(inference):
    """The tension the model should see: predicted and confirmed only.

    - calibration_delta is withheld deliberately. It is predicted minus confirmed,
      a measure of how far the OPERATORS were off — a fact about the instrument,
      not about the felt thought. Sending it invites the model to dramatise the
      pipeline's own error as if it were something the original text contained.
    - Absent stays absent. A missing block reads as unmeasured in the prompt, which
      is what it is; filling it with 0.0 would tell the model the thought was plain.
    """
    t = inference.get("tension")
    if isinstance(t, dict):
        out = {k: t.get(k) for k in ("predicted", "confirmed") if t.get(k) is not None}
        return out or None
    legacy = inference.get("tension_score")
    return {"predicted": legacy} if legacy is not None else None


def call_api(prompt, dry_run=False):
    if dry_run:
        print("[dry-run] prompt:\n")
        print(prompt)
        sys.exit(0)
    # The anthropic SDK needs an API key this box does not have, and it bypassed the
    # privacy gate entirely — reify was the last caller still on that path after
    # vivify.py was migrated off it. llm_call resolves the same capability, dispatches
    # through the claude CLI, enforces the gate, and strips markdown fences.
    # sensitive=False keeps the existing behaviour: these domain voices become the
    # public descriptions in discovery.jsonld, so they are published material already.
    return llm_call(prompt, capability="prose_reconstruction", sensitive=False)


def reify_single(inference, dry_run=False):
    """Reconstruct prose from a single inference's structure."""
    payload = {
        "left_keywords": inference.get("left_keywords", []),
        "clumps": inference.get("clumps", {}),
        "category_paths": inference.get("category_paths", [])[:4],
        "tension": _tension_for_prompt(inference)
    }
    prompt = REIFY_PROMPT + json.dumps(payload, indent=2)
    return call_api(prompt, dry_run=dry_run)


def reify_synthesize(inf_a, inf_b, dry_run=False):
    """Generate text holding two inferences simultaneously."""
    def slim(inf):
        return {
            "left_keywords": inf.get("left_keywords", []),
            "clumps": inf.get("clumps", {}),
            "tension": _tension_for_prompt(inf)
        }
    prompt = SYNTHESIZE_PROMPT.format(
        inf_a=json.dumps(slim(inf_a), indent=2),
        inf_b=json.dumps(slim(inf_b), indent=2)
    )
    return call_api(prompt, dry_run=dry_run)


def reify_voice(category, inferences_dir="inferences", dry_run=False):
    """Generate text that speaks for an entire category, read from the index."""
    index = read_json(Path(inferences_dir) / "index.json")
    ids = ids_for_path(index, category)
    if not ids:
        raise ValueError(f"No inferences indexed under category: {category}")

    inferences = []
    for inf_id in ids:
        inf = load_by_id(inferences_dir, inf_id)
        if inf:
            inferences.append({
                "id": inf["id"],
                "left_keywords": inf.get("left_keywords", []),
                "clumps": inf.get("clumps", {}),
                "tension": _tension_for_prompt(inf)
            })

    if not inferences:
        raise ValueError(f"Indexed ids for {category} not found in storage")

    prompt = VOICE_PROMPT.format(
        category=category,
        inferences=json.dumps(inferences, indent=2)
    )
    return call_api(prompt, dry_run=dry_run)


def reify_domain_voice(inferences_dir, domain_name=None, dry_run=False):
    """Generate one voice for an entire domain — synthesizes all its inferences.

    Unlike reify_voice (a single category path), this speaks for the whole domain
    by loading every inference directly under the domain dir. Used to populate the
    domain descriptions in the JSON-LD discovery surface.
    """
    d = Path(inferences_dir)
    domain_name = domain_name or d.name

    inferences = []
    # rglob, not glob: only `field` stores its inferences flat. logos, pillars and
    # claude_code_sessions use seed/sub nesting, so a direct-children glob found
    # nothing and raised "No inferences found" for three of the four domains — which
    # is why no index has ever carried a voice for them. Same bug, and same fix, as
    # build_index.discover_domains (77f0d2e).
    for path in d.rglob("inf_*.json"):
        inf = read_json(path)
        if inf:
            inferences.append({
                "id": inf["id"],
                "left_keywords": inf.get("left_keywords", []),
                "clumps": inf.get("clumps", {}),
                "tension": _tension_for_prompt(inf)
            })

    if not inferences:
        raise ValueError(f"No inferences found in {d}")

    prompt = VOICE_PROMPT.format(
        category=domain_name,
        inferences=json.dumps(inferences, indent=2)
    )
    return call_api(prompt, dry_run=dry_run)


def usage():
    print(__doc__)
    sys.exit(1)


def main():
    parser = argparse.ArgumentParser(add_help=False)
    parser.add_argument("paths", nargs="*")
    parser.add_argument("--synthesize", action="store_true")
    parser.add_argument("--voice", metavar="CATEGORY")
    parser.add_argument("--dir", default="inferences")
    parser.add_argument("--dry-run", action="store_true")
    parser.add_argument("-h", "--help", action="store_true")
    args = parser.parse_args()

    if args.help:
        usage()

    dry_run = args.dry_run

    if args.voice:
        text = reify_voice(args.voice, inferences_dir=args.dir, dry_run=dry_run)
        print(f"[voice: {args.voice}]\n")
        print(text)
        return

    if args.synthesize:
        if len(args.paths) != 2:
            print("Error: --synthesize requires exactly two inference paths.")
            usage()
        inf_a = read_json(args.paths[0])
        inf_b = read_json(args.paths[1])
        if not inf_a or not inf_b:
            print("Error: could not read one or both inference files.")
            sys.exit(1)
        text = reify_synthesize(inf_a, inf_b, dry_run=dry_run)
        print(f"[synthesis: {inf_a['id']} + {inf_b['id']}]\n")
        print(text)
        return

    if len(args.paths) == 1:
        inf = read_json(args.paths[0])
        if not inf:
            print(f"Error: could not read {args.paths[0]}")
            sys.exit(1)
        text = reify_single(inf, dry_run=dry_run)
        print(f"[reify: {inf['id']}  tension: {_tension_for_prompt(inf) or 'unmeasured'}]\n")
        print(text)
        return

    usage()


if __name__ == "__main__":
    main()

# llm: claude-sonnet-4-6 | 2026-04-17 | repos/vivify-inferences/reify.py | created — inverse pass, JSON inference → prose via Claude API; single/synthesize/voice modes
# llm: claude-sonnet-4-6 | 2026-04-21 | repos/vivify-inferences/reify.py | updated all three prompts — output format changed to series of sentence+bullets constructs
# llm: claude-sonnet-4-6 | 2026-04-27 | repos/vivify-inferences/reify.py | replaced hardcoded model string with resolve_model("prose_reconstruction")
# llm: claude-opus-4-8 | 2026-06-28 | repos/vivify-operators/reify.py | ported index-based reify_voice + added reify_domain_voice (kept operators anthropic-SDK call_api)

# llm: claude-opus-5 | 2026-09-15 | repos/vivify-operators/reify.py | reify_domain_voice rglobs the domain: a direct-children glob raised "No inferences found" for every nested domain (logos, pillars, claude_code_sessions), so --voices could only ever work for flat `field`

# llm: claude-opus-5 | 2026-09-15 | repos/vivify-operators/reify.py | call_api goes through llm_call (claude CLI + privacy gate) instead of the anthropic SDK, which could not authenticate on this box and bypassed the gate — the last caller left on that path

# llm: claude-opus-5 | 2026-09-19 | repos/vivify-operators/reify.py | the prompt described tension as left/right keyword divergence and told the model to lean into that gap — dead since the 2026-07-13 rewire, and this is an executable instruction, not a doc; now sends predicted/confirmed with their real meaning, withholds calibration_delta as an instrument fact, and makes unmeasured explicit instead of readable as low
