Schema-aware generation
比 system prompt 更强的方式——用 function calling / tool use / JSON mode 让 LLM 结构化输出,server 端再转 Rho format。 准确率比 raw prompt 高 1-2 个数量级。
为什么 schema 比 prompt 强
| 方式 | 准确率 | 复杂度 | 适用 |
|---|---|---|---|
| Raw prompt(Prompt templates) | ~80% | 最低 | trial / 简单场景 |
| JSON mode(强制 valid JSON) | ~95% | 中 | production 简单 |
| Function / tool calling(带 schema 验证) | ~99% | 高 | production 严格 |
| Fine-tuning / few-shot | ~99.5%+ | 最高 | 大规模商业 |
核心原理:用 schema 约束 LLM 的输出空间——LLM 不再"自由写 markdown",而是"填这个 JSON 表"。出错可能性大幅降低。
模式 1:JSON mode(OpenAI / Claude / Gemini 都支持)
让 LLM 输出严格 JSON,然后 server 端转 Rho。
Schema 设计
type RhoBlock =
| { type: "callout"; calloutType: "INFO" | "WARN" | "ZEN"; body: string }
| { type: "layout"; cols: number; cards: { accent: AccentColor; content: string }[] }
| { type: "interact"; namespace?: string; controls: Control[]; computed?: Computed[]; template: Template }
| { type: "tabs"; default?: string; tabs: { title: string; content: string }[] }
| { type: "stepper"; linear?: boolean; steps: { title: string; content: string }[] }
| { type: "modal"; trigger: string; body: string; width?: "small" | "medium" | "large" }
| { type: "timeline"; events: { date: string; title: string; body: string; accent?: AccentColor }[] }
| { type: "annotate"; body: string; annotations: { range: string; color: AccentColor; label: string }[] };
type Control =
| { type: "slider"; name: string; min: number; max: number; initial: number; step: number }
| { type: "input"; name: string; default: string }
| { type: "select"; name: string; options: string[] }
| { type: "toggle"; name: string; default: boolean }
| { type: "button"; name: string; label: string; value: string };
type Computed = { name: string; expression: string };
type Template =
| { format: "stl"; lines: string[] }
| { format: "vega-lite"; spec: object }
| { format: "svg"; markup: string };
type AccentColor = "blue" | "red" | "green" | "yellow" | "purple" | "gray";
interface RhoDocument {
title?: string;
blocks: RhoBlock[];
}
Prompt 配 JSON mode
// OpenAI
const response = await openai.chat.completions.create({
model: "gpt-4-turbo",
response_format: { type: "json_object" },
messages: [
{ role: "system", content: "You output a JSON document conforming to RhoDocument schema. Schema:\n```typescript\n" + RhoSchemaTypeDef + "\n```" },
{ role: "user", content: "Generate a BMI calculator." }
],
});
const rhoDoc: RhoDocument = JSON.parse(response.choices[0].message.content!);
const markdown = renderRhoFromJSON(rhoDoc); // 你写的转换器
Server 端转 Rho markdown
function renderRhoFromJSON(doc: RhoDocument): string {
let md = doc.title ? `# ${doc.title}\n\n` : '';
for (const block of doc.blocks) {
md += renderBlock(block) + '\n\n';
}
return md;
}
function renderBlock(block: RhoBlock): string {
switch (block.type) {
case 'callout':
return `> [!${block.calloutType}]\n` + block.body.split('\n').map(l => `> ${l}`).join('\n');
case 'layout':
const cards = block.cards.map(c =>
`:::card accent=${c.accent}\n${c.content}\n:::`
).join('\n\n');
return `\`\`\`layout grid cols=${block.cols}\n${cards}\n\`\`\``;
case 'interact':
return renderInteract(block);
// ... etc.
}
}
优点:LLM 只需输出 JSON,不操心 markdown 缩进 / fenced code 嵌套。
模式 2:Function / tool calling
更精确的版本——告诉 LLM 一个 function schema,让它"call this function with params"。
OpenAI / GPT function calling
const tools = [
{
type: "function",
function: {
name: "create_rho_document",
description: "Create a Rho format markdown document",
parameters: {
type: "object",
properties: {
title: { type: "string" },
blocks: {
type: "array",
items: {
oneOf: [
{
type: "object",
properties: {
type: { const: "callout" },
callout_type: { type: "string", enum: ["INFO", "WARN", "ZEN", "NOTE", "TIP", "IMPORTANT", "WARNING", "CAUTION"] },
body: { type: "string" }
},
required: ["type", "callout_type", "body"]
},
{
type: "object",
properties: {
type: { const: "interact" },
namespace: { type: "string" },
controls: { type: "array", items: { /* ... */ } },
template: { /* ... */ }
},
required: ["type", "controls", "template"]
}
// ... etc.
]
}
}
},
required: ["blocks"]
}
}
}
];
const response = await openai.chat.completions.create({
model: "gpt-4-turbo",
messages: [{ role: "user", content: "BMI calculator with weight slider 40-120 and height slider 1.4-2.1" }],
tools,
tool_choice: { type: "function", function: { name: "create_rho_document" } }
});
const args = JSON.parse(response.choices[0].message.tool_calls![0].function.arguments);
const markdown = renderRhoFromJSON(args);
Claude tool use
const tools = [
{
name: "create_rho_document",
description: "Create a Rho format markdown document",
input_schema: { /* same as above */ }
}
];
const response = await anthropic.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 4096,
tools,
messages: [{ role: "user", content: "BMI calculator..." }]
});
const toolUse = response.content.find(b => b.type === 'tool_use');
const args = toolUse.input;
const markdown = renderRhoFromJSON(args);
Gemini function calling
const model = genAI.getGenerativeModel({
model: "gemini-2.5-pro",
tools: [{ functionDeclarations: [{ name: "create_rho_document", description: "...", parameters: { /* JSON Schema */ } }] }]
});
模式 3:Few-shot examples(最简加强)
不用 schema 也能拉准确率——喂 LLM 5-10 个真实 Rho 文档作 examples:
const fewShot = `
EXAMPLE 1:
User: BMI calculator
Output:
\`\`\`markdown
# BMI Calculator
> [!INFO]
> Formula: BMI = weight / height²
\`\`\`interact
slider weight 40 120 70 1
slider height 1.4 2.1 1.7 0.01
computed bmi = weight / (height * height)
template stl:
[BMI] -> [{bmi:.1f}]
\`\`\`
\`\`\`
EXAMPLE 2:
User: Compound interest
Output: ...
EXAMPLE 3:
User: Pendulum animation
Output: ...
`;
const response = await openai.chat.completions.create({
model: "gpt-4-turbo",
messages: [
{ role: "system", content: "You generate Rho format. Examples below.\n\n" + fewShot },
{ role: "user", content: "Create a mortgage calculator" }
]
});
示例越多越好——5 个例子 vs 0 个例子,准确率涨 ~10-15 个百分点。
模式 4:组合(推荐 production)
LLM Protocol (system message, ~50 lines)
+
5-10 Few-shot examples(覆盖所有 capability)
+
Function calling / JSON mode(结构化输出)
+
Server 端 schema validation(拒绝非法输出)
+
Fallback:解析失败再 prompt 一次(带错误 hint)
这种组合实测准确率 99%+。
Server 端验证
LLM 即使 schema-driven 也可能出错(比如 mini DSL 表达式语法错)。你应该跑一次解析校验:
import { unified } from 'unified';
import remarkParse from 'remark-parse';
import { remarkPlugins } from '@rho/md';
const processor = unified().use(remarkParse);
remarkPlugins.forEach(p => processor.use(p));
async function validateRho(markdown: string): Promise<{ valid: boolean; errors: string[] }> {
try {
const ast = await processor.parse(markdown);
// 自定义校验:每个 interact 块的 mini DSL 表达式 / vega-lite spec / etc.
const errors = customValidate(ast);
return { valid: errors.length === 0, errors };
} catch (e) {
return { valid: false, errors: [String(e)] };
}
}
// 失败回写:
const result = await validateRho(llmOutput);
if (!result.valid) {
// 重新 prompt:告诉 LLM 错在哪里
const retry = await llm.generate({
messages: [
...originalMessages,
{ role: "assistant", content: llmOutput },
{ role: "user", content: `Errors: ${result.errors.join(', ')}. Please fix.` }
]
});
}
Performance / cost
| 模式 | 每次调用 token 量 | 准确率 | 适用规模 |
|---|---|---|---|
| Raw prompt | ~500-1000 | 80% | 小,免费试 |
| Few-shot | ~1500-3000 | 90% | 中等 |
| Function calling + few-shot | ~2000-4000 | 99% | production |
| Fine-tune | ~500-1000 (短 prompt) | 99.5% | 大规模 |
Token cost 跟准确率正相关——production 一般用 function calling 比 raw prompt 总成本反而低(少 retry / 少 fallback)。
See also
- Prompt templates — 简单 prompt 起步
- Common LLM mistakes & repair — 解析失败时的修复 patterns
- AINP 协议 — §21 单页速查表 — 完整 schema 参考源
- AINP 协议 — §16 给 LLM 作者
- Developer Reference: API — 用
@rho/md校验 LLM 输出