Schema-aware generation

Stronger than system prompts — use function calling / tool use / JSON mode to make the LLM produce structured output; convert to Rho format on the server. Accuracy 1-2 orders of magnitude higher than raw prompts.


Why schema beats prompt

Mode Accuracy Complexity Use
Raw prompt (Prompt templates) ~80% Lowest Trial / simple
JSON mode (forces valid JSON) ~95% Medium Production-simple
Function / tool calling (with schema validation) ~99% High Production-strict
Fine-tuning / few-shot ~99.5%+ Highest Large-scale commercial

Core principle: schema constrains the LLM's output space — instead of "freely write markdown," it's "fill out this JSON form." Error chance drops dramatically.


Mode 1: JSON mode (OpenAI / Claude / Gemini all support)

Have the LLM output strict JSON; convert to Rho on the server.

Schema design

type RhoBlock =
  | { type: "callout"; calloutType: "INFO" | "WARN" | "ZEN"; body: string }
  | { type: "layout"; cols: number; cards: { accent: AccentColor; content: string }[] }
  | { type: "interact"; namespace?: string; controls: Control[]; computed?: Computed[]; template: Template }
  | { type: "tabs"; default?: string; tabs: { title: string; content: string }[] }
  | { type: "stepper"; linear?: boolean; steps: { title: string; content: string }[] }
  | { type: "modal"; trigger: string; body: string; width?: "small" | "medium" | "large" }
  | { type: "timeline"; events: { date: string; title: string; body: string; accent?: AccentColor }[] }
  | { type: "annotate"; body: string; annotations: { range: string; color: AccentColor; label: string }[] };

type Control =
  | { type: "slider"; name: string; min: number; max: number; initial: number; step: number }
  | { type: "input"; name: string; default: string }
  | { type: "select"; name: string; options: string[] }
  | { type: "toggle"; name: string; default: boolean }
  | { type: "button"; name: string; label: string; value: string };

type Computed = { name: string; expression: string };

type Template =
  | { format: "stl"; lines: string[] }
  | { format: "vega-lite"; spec: object }
  | { format: "svg"; markup: string };

type AccentColor = "blue" | "red" | "green" | "yellow" | "purple" | "gray";

interface RhoDocument {
  title?: string;
  blocks: RhoBlock[];
}

Prompt with JSON mode

// OpenAI
const response = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  response_format: { type: "json_object" },
  messages: [
    { role: "system", content: "You output a JSON document conforming to RhoDocument schema. Schema:\n```typescript\n" + RhoSchemaTypeDef + "\n```" },
    { role: "user", content: "Generate a BMI calculator." }
  ],
});

const rhoDoc: RhoDocument = JSON.parse(response.choices[0].message.content!);
const markdown = renderRhoFromJSON(rhoDoc);  // Your conversion

Server-side conversion to Rho markdown

function renderRhoFromJSON(doc: RhoDocument): string {
  let md = doc.title ? `# ${doc.title}\n\n` : '';
  for (const block of doc.blocks) {
    md += renderBlock(block) + '\n\n';
  }
  return md;
}

function renderBlock(block: RhoBlock): string {
  switch (block.type) {
    case 'callout':
      return `> [!${block.calloutType}]\n` + block.body.split('\n').map(l => `> ${l}`).join('\n');

    case 'layout':
      const cards = block.cards.map(c =>
        `:::card accent=${c.accent}\n${c.content}\n:::`
      ).join('\n\n');
      return `\`\`\`layout grid cols=${block.cols}\n${cards}\n\`\`\``;

    case 'interact':
      return renderInteract(block);

    // ... etc.
  }
}

Advantage: the LLM only outputs JSON; no need to worry about markdown indentation / fenced code nesting.


Mode 2: Function / tool calling

More precise — give the LLM a function schema; have it "call this function with params."

OpenAI / GPT function calling

const tools = [
  {
    type: "function",
    function: {
      name: "create_rho_document",
      description: "Create a Rho format markdown document",
      parameters: {
        type: "object",
        properties: {
          title: { type: "string" },
          blocks: {
            type: "array",
            items: {
              oneOf: [
                {
                  type: "object",
                  properties: {
                    type: { const: "callout" },
                    callout_type: { type: "string", enum: ["INFO", "WARN", "ZEN", "NOTE", "TIP", "IMPORTANT", "WARNING", "CAUTION"] },
                    body: { type: "string" }
                  },
                  required: ["type", "callout_type", "body"]
                },
                {
                  type: "object",
                  properties: {
                    type: { const: "interact" },
                    namespace: { type: "string" },
                    controls: { type: "array", items: { /* ... */ } },
                    template: { /* ... */ }
                  },
                  required: ["type", "controls", "template"]
                }
                // ... etc.
              ]
            }
          }
        },
        required: ["blocks"]
      }
    }
  }
];

const response = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  messages: [{ role: "user", content: "BMI calculator with weight slider 40-120 and height slider 1.4-2.1" }],
  tools,
  tool_choice: { type: "function", function: { name: "create_rho_document" } }
});

const args = JSON.parse(response.choices[0].message.tool_calls![0].function.arguments);
const markdown = renderRhoFromJSON(args);

Claude tool use

const tools = [
  {
    name: "create_rho_document",
    description: "Create a Rho format markdown document",
    input_schema: { /* same as above */ }
  }
];

const response = await anthropic.messages.create({
  model: "claude-sonnet-4-6",
  max_tokens: 4096,
  tools,
  messages: [{ role: "user", content: "BMI calculator..." }]
});

const toolUse = response.content.find(b => b.type === 'tool_use');
const args = toolUse.input;
const markdown = renderRhoFromJSON(args);

Gemini function calling

const model = genAI.getGenerativeModel({
  model: "gemini-2.5-pro",
  tools: [{ functionDeclarations: [{ name: "create_rho_document", description: "...", parameters: { /* JSON Schema */ } }] }]
});

Mode 3: Few-shot examples (simplest boost)

You can lift accuracy without schema — feed the LLM 5-10 real Rho documents as examples:

const fewShot = `
EXAMPLE 1:
User: BMI calculator
Output:
\`\`\`markdown
# BMI Calculator

> [!INFO]
> Formula: BMI = weight / height²

\`\`\`interact
slider weight 40 120 70 1
slider height 1.4 2.1 1.7 0.01
computed bmi = weight / (height * height)
template stl:
[BMI] -> [{bmi:.1f}]
\`\`\`
\`\`\`

EXAMPLE 2:
User: Compound interest
Output: ...

EXAMPLE 3:
User: Pendulum animation
Output: ...
`;

const response = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  messages: [
    { role: "system", content: "You generate Rho format. Examples below.\n\n" + fewShot },
    { role: "user", content: "Create a mortgage calculator" }
  ]
});

More examples = better — 5 examples vs 0 examples lifts accuracy by ~10-15 percentage points.


LLM Protocol (system message, ~50 lines)
    +
5-10 few-shot examples (covering all capabilities)
    +
Function calling / JSON mode (structured output)
    +
Server-side schema validation (reject invalid output)
    +
Fallback: re-prompt on parse failure (with error hint)

This combo measures 99%+ accuracy.


Server-side validation

LLMs even when schema-driven can still fail (e.g., mini DSL expression syntax error). You should run a parse check:

import { unified } from 'unified';
import remarkParse from 'remark-parse';
import { remarkPlugins } from '@rho/md';

const processor = unified().use(remarkParse);
remarkPlugins.forEach(p => processor.use(p));

async function validateRho(markdown: string): Promise<{ valid: boolean; errors: string[] }> {
  try {
    const ast = await processor.parse(markdown);
    // Custom validation: each interact block's mini DSL / vega-lite spec / etc.
    const errors = customValidate(ast);
    return { valid: errors.length === 0, errors };
  } catch (e) {
    return { valid: false, errors: [String(e)] };
  }
}

// Failure feedback:
const result = await validateRho(llmOutput);
if (!result.valid) {
  // Re-prompt: tell the LLM what went wrong
  const retry = await llm.generate({
    messages: [
      ...originalMessages,
      { role: "assistant", content: llmOutput },
      { role: "user", content: `Errors: ${result.errors.join(', ')}. Please fix.` }
    ]
  });
}

Performance / cost

Mode Tokens per call Accuracy Suitable for
Raw prompt ~500-1000 80% Small, free trial
Few-shot ~1500-3000 90% Mid-size
Function calling + few-shot ~2000-4000 99% Production
Fine-tune ~500-1000 (short prompt) 99.5% Large-scale

Token cost correlates with accuracy — production typically finds function calling lower total cost than raw prompt (fewer retries / fewer fallbacks).


See also