Schema-aware generation

比 system prompt 更强的方式——用 function calling / tool use / JSON mode 让 LLM 结构化输出,server 端再转 Rho format。 准确率比 raw prompt 高 1-2 个数量级。


为什么 schema 比 prompt 强

方式 准确率 复杂度 适用
Raw prompt(Prompt templates) ~80% 最低 trial / 简单场景
JSON mode(强制 valid JSON) ~95% 中 production 简单
Function / tool calling(带 schema 验证) ~99% 高 production 严格
Fine-tuning / few-shot ~99.5%+ 最高 大规模商业

核心原理:用 schema 约束 LLM 的输出空间——LLM 不再"自由写 markdown",而是"填这个 JSON 表"。出错可能性大幅降低。


模式 1:JSON mode(OpenAI / Claude / Gemini 都支持)

让 LLM 输出严格 JSON,然后 server 端转 Rho。

Schema 设计

type RhoBlock =
  | { type: "callout"; calloutType: "INFO" | "WARN" | "ZEN"; body: string }
  | { type: "layout"; cols: number; cards: { accent: AccentColor; content: string }[] }
  | { type: "interact"; namespace?: string; controls: Control[]; computed?: Computed[]; template: Template }
  | { type: "tabs"; default?: string; tabs: { title: string; content: string }[] }
  | { type: "stepper"; linear?: boolean; steps: { title: string; content: string }[] }
  | { type: "modal"; trigger: string; body: string; width?: "small" | "medium" | "large" }
  | { type: "timeline"; events: { date: string; title: string; body: string; accent?: AccentColor }[] }
  | { type: "annotate"; body: string; annotations: { range: string; color: AccentColor; label: string }[] };

type Control =
  | { type: "slider"; name: string; min: number; max: number; initial: number; step: number }
  | { type: "input"; name: string; default: string }
  | { type: "select"; name: string; options: string[] }
  | { type: "toggle"; name: string; default: boolean }
  | { type: "button"; name: string; label: string; value: string };

type Computed = { name: string; expression: string };

type Template =
  | { format: "stl"; lines: string[] }
  | { format: "vega-lite"; spec: object }
  | { format: "svg"; markup: string };

type AccentColor = "blue" | "red" | "green" | "yellow" | "purple" | "gray";

interface RhoDocument {
  title?: string;
  blocks: RhoBlock[];
}

Prompt 配 JSON mode

// OpenAI
const response = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  response_format: { type: "json_object" },
  messages: [
    { role: "system", content: "You output a JSON document conforming to RhoDocument schema. Schema:\n```typescript\n" + RhoSchemaTypeDef + "\n```" },
    { role: "user", content: "Generate a BMI calculator." }
  ],
});

const rhoDoc: RhoDocument = JSON.parse(response.choices[0].message.content!);
const markdown = renderRhoFromJSON(rhoDoc);  // 你写的转换器

Server 端转 Rho markdown

function renderRhoFromJSON(doc: RhoDocument): string {
  let md = doc.title ? `# ${doc.title}\n\n` : '';
  for (const block of doc.blocks) {
    md += renderBlock(block) + '\n\n';
  }
  return md;
}

function renderBlock(block: RhoBlock): string {
  switch (block.type) {
    case 'callout':
      return `> [!${block.calloutType}]\n` + block.body.split('\n').map(l => `> ${l}`).join('\n');

    case 'layout':
      const cards = block.cards.map(c =>
        `:::card accent=${c.accent}\n${c.content}\n:::`
      ).join('\n\n');
      return `\`\`\`layout grid cols=${block.cols}\n${cards}\n\`\`\``;

    case 'interact':
      return renderInteract(block);

    // ... etc.
  }
}

优点:LLM 只需输出 JSON,不操心 markdown 缩进 / fenced code 嵌套。


模式 2:Function / tool calling

更精确的版本——告诉 LLM 一个 function schema,让它"call this function with params"。

OpenAI / GPT function calling

const tools = [
  {
    type: "function",
    function: {
      name: "create_rho_document",
      description: "Create a Rho format markdown document",
      parameters: {
        type: "object",
        properties: {
          title: { type: "string" },
          blocks: {
            type: "array",
            items: {
              oneOf: [
                {
                  type: "object",
                  properties: {
                    type: { const: "callout" },
                    callout_type: { type: "string", enum: ["INFO", "WARN", "ZEN", "NOTE", "TIP", "IMPORTANT", "WARNING", "CAUTION"] },
                    body: { type: "string" }
                  },
                  required: ["type", "callout_type", "body"]
                },
                {
                  type: "object",
                  properties: {
                    type: { const: "interact" },
                    namespace: { type: "string" },
                    controls: { type: "array", items: { /* ... */ } },
                    template: { /* ... */ }
                  },
                  required: ["type", "controls", "template"]
                }
                // ... etc.
              ]
            }
          }
        },
        required: ["blocks"]
      }
    }
  }
];

const response = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  messages: [{ role: "user", content: "BMI calculator with weight slider 40-120 and height slider 1.4-2.1" }],
  tools,
  tool_choice: { type: "function", function: { name: "create_rho_document" } }
});

const args = JSON.parse(response.choices[0].message.tool_calls![0].function.arguments);
const markdown = renderRhoFromJSON(args);

Claude tool use

const tools = [
  {
    name: "create_rho_document",
    description: "Create a Rho format markdown document",
    input_schema: { /* same as above */ }
  }
];

const response = await anthropic.messages.create({
  model: "claude-sonnet-4-6",
  max_tokens: 4096,
  tools,
  messages: [{ role: "user", content: "BMI calculator..." }]
});

const toolUse = response.content.find(b => b.type === 'tool_use');
const args = toolUse.input;
const markdown = renderRhoFromJSON(args);

Gemini function calling

const model = genAI.getGenerativeModel({
  model: "gemini-2.5-pro",
  tools: [{ functionDeclarations: [{ name: "create_rho_document", description: "...", parameters: { /* JSON Schema */ } }] }]
});

模式 3:Few-shot examples(最简加强)

不用 schema 也能拉准确率——喂 LLM 5-10 个真实 Rho 文档作 examples:

const fewShot = `
EXAMPLE 1:
User: BMI calculator
Output:
\`\`\`markdown
# BMI Calculator

> [!INFO]
> Formula: BMI = weight / height²

\`\`\`interact
slider weight 40 120 70 1
slider height 1.4 2.1 1.7 0.01
computed bmi = weight / (height * height)
template stl:
[BMI] -> [{bmi:.1f}]
\`\`\`
\`\`\`

EXAMPLE 2:
User: Compound interest
Output: ...

EXAMPLE 3:
User: Pendulum animation
Output: ...
`;

const response = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  messages: [
    { role: "system", content: "You generate Rho format. Examples below.\n\n" + fewShot },
    { role: "user", content: "Create a mortgage calculator" }
  ]
});

示例越多越好——5 个例子 vs 0 个例子,准确率涨 ~10-15 个百分点。


模式 4:组合(推荐 production)

LLM Protocol (system message, ~50 lines)
    +
5-10 Few-shot examples(覆盖所有 capability)
    +
Function calling / JSON mode(结构化输出)
    +
Server 端 schema validation(拒绝非法输出)
    +
Fallback:解析失败再 prompt 一次(带错误 hint)

这种组合实测准确率 99%+。


Server 端验证

LLM 即使 schema-driven 也可能出错(比如 mini DSL 表达式语法错)。你应该跑一次解析校验:

import { unified } from 'unified';
import remarkParse from 'remark-parse';
import { remarkPlugins } from '@rho/md';

const processor = unified().use(remarkParse);
remarkPlugins.forEach(p => processor.use(p));

async function validateRho(markdown: string): Promise<{ valid: boolean; errors: string[] }> {
  try {
    const ast = await processor.parse(markdown);
    // 自定义校验:每个 interact 块的 mini DSL 表达式 / vega-lite spec / etc.
    const errors = customValidate(ast);
    return { valid: errors.length === 0, errors };
  } catch (e) {
    return { valid: false, errors: [String(e)] };
  }
}

// 失败回写:
const result = await validateRho(llmOutput);
if (!result.valid) {
  // 重新 prompt:告诉 LLM 错在哪里
  const retry = await llm.generate({
    messages: [
      ...originalMessages,
      { role: "assistant", content: llmOutput },
      { role: "user", content: `Errors: ${result.errors.join(', ')}. Please fix.` }
    ]
  });
}

Performance / cost

模式 每次调用 token 量 准确率 适用规模
Raw prompt ~500-1000 80% 小,免费试
Few-shot ~1500-3000 90% 中等
Function calling + few-shot ~2000-4000 99% production
Fine-tune ~500-1000 (短 prompt) 99.5% 大规模

Token cost 跟准确率正相关——production 一般用 function calling 比 raw prompt 总成本反而低(少 retry / 少 fallback)。


See also