Architecture Migration · MVP Blueprint

Video2SOP → Dify Workflow

多 Agent 视频转 SOP 系统迁移至 Dify 平台的 MVP 工作流设计

Source: video2sop (6,200 LOC) Target: Dify-compatible Agent Platform Scope: MVP (5/7 stages migratable) 2026-07-13
5
可迁移 LLM/Code 节点
3
外部微服务 API
2
不可迁移 (实时/人工)
~85%
核心功能覆盖

MVP Workflow Chart

▶
Start Node
视频输入
video_url + granularity (small / medium / large)
video_url
⇅
HTTP Request · External API
音频提取 + ASR 转录
调用外部微服务:FFmpeg 提音 → Whisper / 火山引擎 ASR → 转录稿 JSON
transcript_json
{ }
Code Node · Python
转录稿清洗
压缩 word-level timestamp → compact text。纯文本处理,100 行直搬
compact_transcript
✦
LLM Node · JSON Output
Orchestrator 任务分析
分析转录稿 → 分类领域 (repair/production/testing/training) → 输出 TaskPlan JSON
task_plan (JSON)
✦
LLM Node · JSON Output
Process Structurer
核心节点:转录稿 → 结构化步骤 JSON(step_number, phase, capture_strategy, description, timestamps, confidence)
structured_data (JSON)
⇅
HTTP Request · External API
帧提取 + 模糊评分
调用外部微服务:FFmpeg 抽帧 → OpenCV Laplacian 评分 → 返回最优截图 URL
frame_results + image_urls
{ }
Code Node · Python + IF/ELSE
QA 校验
检查截图缺失 / 低置信度步骤 → 输出 qa_report(issues[], low_confidence[])
qa_report
T
Template Node · Jinja2
Markdown SOP 渲染
structured_data + frame_results → 替换式 SOP Markdown 文档
sop_markdown
■
End Node
SOP 输出
sop_markdown + qa_report + structured_data + image_urls

需保留为外部微服务(3 个 API)

Dify 无 FFmpeg / OpenCV / 本地模型推理能力,以下模块封装为 HTTP API,由 Dify HTTP Request 节点调用。

POST
/api/transcribe
视频 → 音频提取 → ASR 转录 → 转录稿清洗。一步到位返回 compact_transcript。
← audio_extractor.py + transcriber.py + transcript_cleaner.py
POST
/api/extract-frames
视频 + structured_data → 按步骤时间窗抽帧 → Laplacian 评分 → 返回最优截图 + 候选图。
← frame_extractor.py (660 LOC)
POST
/api/render
如果 Dify Template 节点不足以渲染复杂 SOP,用 Python 兜底渲染。
← markdown_renderer.py (602 LOC) · 可选

迁移可行性矩阵

现有组件 LOC Dify 节点 难度 状态 说明
Orchestrator 任务分析 345 LLM Node (JSON) 低 可迁移 system_prompt 原样搬,response_format Dify 原生支持
LLM Structurer 2,264 LLM Node (JSON) 中 可迁移 核心是 prompt 工程,粒度控制用变量注入
Transcript Cleaner 100 Code Node (Python) 低 可迁移 纯文本处理,直搬
QA Inspector ~80 Code Node + IF/ELSE 低 可迁移 校验逻辑 + 条件分支
Markdown Renderer 602 Template (Jinja2) 中 需精简 602 行渲染逻辑需精简为 Jinja2 模板,复杂部分可走外部 API
Audio Extract + ASR 543 HTTP Request 中 外部 API FFmpeg + Whisper/火山,封装为 /api/transcribe
Frame Extractor 660 HTTP Request 中 外部 API FFmpeg + OpenCV,封装为 /api/extract-frames
EventBus 实时事件流 605 — — 不可迁移 Dify 工作流是同步的,不支持 SSE 事件广播
Agent 间消息通信 ~200 — — 不可迁移 Dify Agent 孤立,不支持 Agent-to-Agent 定向消息
Human-in-the-loop 截图审核 145 — — 需拆分 Dify 不支持中途暂停,拆为独立的审核工作流
断点续跑 (--skip-to) ~100 — — 不可迁移 Dify 工作流不支持中间阶段恢复

可直接搬运的 Prompt 资产

以下 Prompt 已在生产环境验证,迁移到 Dify LLM 节点时只需配置 system_prompt + user_message 模板。

① Orchestrator System Prompt

来源:orchestrator_agent.py → ORCHESTRATOR_SYSTEM_PROMPT

输入:transcript_preview (前 3000 字符) + video_metadata

输出 JSON:domain, task_summary, complexity, estimated_steps, agents_required, documents_to_produce, reasoning

system_prompt · orchestrator
Role: Orchestrator - task analysis agent in multi-agent SOP system

Input: transcript preview (first ~3000 chars) + video metadata

Output JSON:
{
  "domain": "repair|production|testing|training|mixed",
  "task_summary": "One concise sentence",
  "complexity": "low|medium|high",
  "estimated_steps": 8,
  "agents_required": ["repair"],
  "documents_to_produce": [{"type":"replacement_procedure","agent":"repair"}],
  "reasoning": "1-2 sentences"
}

Domain cues:
  repair     → "remove","replace","screw","拆卸","更换","维修"
  production → "install","assemble","装配","产线","工位"
  testing    → "test","measure","verify","测试","检验"
  training   → "learn","practice","学习","培训","教学"
② Process Structurer System Prompt

来源:prompts/system_prompt.txt (114 行完整版)

输入:compact_transcript + granularity (small/medium/large)

输出 JSON:title + steps[],每步含 step_number, phase, capture_strategy, step_name, description, extract_start/end_sec, preferred_timestamp_sec, confidence, confidence_reason

system_prompt · structurer (核心规则摘录)
Task: Convert word-level timestamped transcript → structured JSON

Step fields (10 per step):
  step_number      : int
  phase            : preparatory|removal|installation|startup|software|final
  capture_strategy : before|during|after|stable|stable_ui
  step_name        : max 8 words, imperative
  description      : professional English, imperative voice
  extract_start_sec: float (1s before action verb)
  extract_end_sec  : float (2-3s after action verb)
  preferred_timestamp_sec: float (best screenshot moment)
  confidence       : 0.0-1.0
  confidence_reason: one sentence explanation

Granularity profiles:
  small  → 12-18 steps, high detail, split micro-actions
  medium → 8-12 steps, balanced
  large  → 5-8 steps, concise, merge related actions

Critical rules:
  - Strict chronological order (timestamps always increasing)
  - Screenshot window MUST match the step's own transcript evidence
  - OUTPUT ONLY VALID JSON

架构降级与权衡

维度 现有 Python 项目 Dify MVP 工作流 影响
多 Agent 协作 真多 Agent,EventBus + 消息通信 线性工作流,Agent 间无对话 降级
实时可观测性 War Room SSE 实时展示 Agent 协作 Dify 执行日志 降级
人工审核环节 image_review.md + 断点续跑 需拆为独立的审核工作流 需拆分
LLM Prompt 管理 代码内硬编码 Dify 可视化管理 + 版本控制 提升
模型切换 改 config.py + 重启 Dify UI 一键切换 提升
部署与分享 自建 Web 服务 平台托管,团队直接用 提升
可扩展性 改代码 + 测试 拖拽节点 提升

MVP 关键决策

什么进了 MVP
  • 5 个 Dify 原生节点:2× LLM (Orchestrator + Structurer) + 2× Code (Clean + QA) + 1× Template (Render)
  • 2 个 HTTP Request 节点:调用外部 ASR + 帧提取微服务
  • 2 套已验证的 Prompt:直接从生产代码中提取,含 JSON Schema 和粒度控制
  • 1 个 Start + 1 个 End:标准化输入输出
什么没进 MVP
  • War Room 实时事件流:Dify 架构不支持 SSE 广播,牺牲实时可观测性
  • Agent 间消息通信:Dify Agent 孤立,多 Agent 协作降级为线性流水线
  • Human-in-the-loop:截图审核需拆为第二个工作流,MVP 先做自动生成
  • 断点续跑:Dify 不支持中间阶段恢复,每次完整执行
MVP 之后可以做什么
  • 审核工作流:独立 Dify Workflow,输入 structured_data → 人工修改 → 重新渲染
  • 多文档类型:Orchestrator 输出的 documents_to_produce 驱动不同渲染模板
  • 知识库 RAG:Dify 原生支持知识库,可挂载历史 SOP 作为参考
  • 多模态帧分析:用 LLM Vision 节点替代 OpenCV 评分,提升截图质量