Back to Learning Matrix
M2 Say goodbye to cumbersome system interfaces and complex operations — speak to query data, handle business, and control devices

Multimodal AI Interaction

Based on physical AI terminals, integrating edge vision, voice, and business-system APIs to enable multimodal spatial interaction.

Multimodal AI Interaction Illustration

Core Hardware

  • WATCHER SenseCAP Watcher XiaoZhi English Edition On-Device Multimodal Interaction Terminal · Voice Capture & Vision Recognition Entry
  • GW reComputer RK3588-40 Edge Intelligence Controller · Business System & MCP Bridge Host
  • HMI reTerminal D1001 8-Inch Industrial Smart Touchscreen · Workstation HMI
  • GPU reComputer J4012 (Jetson Orin NX 16GB) L3 Advanced Compute Host · Pure Local Offline Voice Pipeline

What This Module Solves

In scenarios such as warehouse management, exhibition hall guidance, and smart reception desks, on-site staff must stop manual operations and retrieve business data via keyboard or mobile phone, resulting in low efficiency. Traditional interaction terminals lack visual context and cannot proactively detect approaching personnel or abnormal actions. Smart terminals are often closed ecosystems, making integration with existing WMS/ERP/CRM systems difficult; some industrial and government/enterprise scenarios prohibit audio and business data from being uploaded to the public internet.

Difficulty
Advanced
Duration
L1 1 day / L2 2–3 days / L3 3–5 days
Shortest Format
1 day (Taster Session · L1)
Teaching Format
3 tiers: Taster / Workshop / Bootcamp
Core Protocols
MCP / REST API / Wi-Fi
L3 Offline Compute
100 TOPS(Jetson Orin NX 16GB)

L1: zero baseline or first-time exposure to edge AI interaction devices. L2: Docker fundamentals and REST API calling experience required. L3: Linux, PyTorch/Jetson fundamentals, and shell operation skills required.

Typical Scenarios

Smart warehousing and workshop management: hands-free inventory lookup, voice-based stock-in/out logging, visual alerts for anomalous materials Exhibition halls and public guidance: proactive visitor detection, multilingual voice narration, booth linkage control Smart reception and meeting spaces: visitor reception, voice control of room devices, automated schedule queries Assisted care and spatial services: specific-behavior detection alerts, voice help requests, and remote linkage Voice interaction in industrial isolated zones: pure LAN offline voice queries and device status announcements

Core Hardware

  • SenseCAP Watcher (XiaoZhi English Edition)
  • reComputer RK3588-40 Edge Intelligence Controller
  • reTerminal D1001 8-Inch Smart Touchscreen
  • reComputer J4012(Jetson Orin NX 16GB)

Key Capabilities

  • On-device vision object detection and event triggering
  • Natural-language voice queries and Agent role configuration
  • Business-system tool calling based on the MCP protocol
  • Local WMS system Docker deployment and API integration
  • Pure local offline voice AI pipeline deployment (VAD→ASR→LLM→TTS)

Course Hardware

This course centers on "visible AI terminals + local inference compute".

  • SenseCAP Watcher XiaoZhi English Edition

    SenseCAP Watcher (XiaoZhi English Edition)

    On-device multimodal interaction terminal, voice capture and vision recognition entry point

    Integrates audio/video capture and screen display, supports object detection, person-approach sensing, and natural-language voice interaction; connects to the SenseCraft AI platform via Wi-Fi. SKU 100051523, 2 units per group.

  • reComputer RK3588-40 Edge Intelligence Controller

    reComputer RK3588-40 Edge Intelligence Controller

    Edge host running business systems and MCP bridge services

    16GB RAM, 6 TOPS compute, running Docker-containerized WMS warehouse system and MCP Bridge service, enabling integration between LAN business data and LLM tool calling. SKU 100086238, includes 12V power adapter.

  • reTerminal D1001 8-Inch Smart Touchscreen

    reTerminal D1001 8-Inch Smart Touchscreen

    Workstation HMI, warehouse business data entry and status monitoring

    8-inch industrial smart touch terminal, with camera and dual microphones, SKU 100058144. Serves as a workstation HMI for warehouse business data entry and status monitoring, can directly connect to the host to display the WMS management console and interaction logs.

  • reComputer J4012 Jetson Orin NX 16GB

    reComputer J4012(Jetson Orin NX 16GB)

    L3 advanced edge compute host, deploying pure local offline voice pipeline

    100 TOPS-class compute, pre-installed JetPack/CUDA/TensorRT/PyTorch environment, deploying Silero VAD + Whisper ASR + Qwen LLM + ChatTTS pure local offline voice AI pipeline, usable even without network. SKU 114110314, with 19V/4.7A power adapter.

另配便携式现场显示器(13.3" 1080P)、CUDY AX3000 Wi-Fi 6路由器、供电排插、六类千兆网线、智能仓管WMS实操模拟物料包(条码标贴/货位标签/实体样本盒)、Watcher桌面支架等通用配件。

Codecraft helps you dare to make, aily-blockly helps you finish it

M0 adopts dual-platform relay toolchain for zero-install, 5-minute results.

SenseCraft AI + Watcher

Cloud LLM + On-Device Multimodal Terminal · Zero-Code Configuration

  1. Watcher Network Setup & Binding
  2. Vision Model & Agent Role Configuration
  3. Voice Query & MCP Tool Calling Experience

MCP Bridge + Docker + OpenClaw

Standard Protocol Bridging · LAN Business Data Stays In-Domain

  1. Local WMS Docker Deployment
  2. MCP Bridge config.yml Configuration
  3. OpenClaw Automation Tool Registration & Integration Testing

Jetson Orin NX + Offline Voice Pipeline

VAD→ASR→LLM→TTS Pure Local Closed Loop · Zero Public-Network Dependency

  1. JetPack Environment Verification
  2. Quantized Model Deployment & VRAM Tuning
  3. Offline Integration Testing & Latency Optimization
Key Turning Point · From Cloud Collaboration to Local Offline Private DeploymentThe SenseCraft AI cloud solution addresses "rapid verification and out-of-box usability"; MCP bridging closes the business data loop within the LAN for the first time, with core inventory and business data staying in-domain; the Jetson offline pipeline completely cuts public-network dependency, achieving zero-external-network voice interaction in high-privacy and industrial isolated-network environments.

L1/L2 relies on internet connection to LLM services; L3 requires 100 TOPS-class edge compute (Jetson Orin NX 16GB), RK3588-40 (6 TOPS) cannot host local LLM inference.

Three-tier Progression: Demo → Consultant → Design

L1 · Demo Level 1 days

Multimodal Interaction Capability Experience

Own a dedicated AI voice assistant, directly query data and control devices through everyday speech

  • Understand the technical architecture combining edge vision and LLM Agent
  • Master the core role of the MCP protocol in integrating on-device AI with business systems
  • Master the selection logic between cloud collaboration and local deployment across business scenarios
L2 · Consultant Level 3 days

Business System Integration and Linkage Configuration

Integrate internal business systems, let voice interaction directly flow work orders and simplify cumbersome operations

  • Independently configure Watcher vision and voice Agent parameters
  • Master Docker-based deployment of local business systems and MCP bridge services
  • Master the method of extending new business APIs via the MCP protocol
L3 · Design Level 5 days

End-to-End Local Offline Voice AI Pipeline Deployment

Achieve pure local offline deployment, usable without network and core business data never leaves the intranet

  • Master the complete local end-to-end VAD→ASR→LLM→TTS voice pipeline architecture
  • Master LLM quantization and deployment optimization methods on Jetson edge computing hardware
  • Capable of delivering AI interaction solutions in high-privacy and industrial isolated-network environments

Curriculum / 15 teaching modules

Same module order, you choose the cut

Select a format to see which modules it covers.

15teaching modules coverage across formats
Module / OutputTaster1 dayWorkshop2–3 daysBootcamp3–5 days
01Pre-class Preparation and Environment Pre-checkHardware bench inventory, network connectivity testing, Watcher firmware pre-check and platform account initialization, RK3588-40 Docker environment and WMS image pre-loading, J4012 JetPack environment pre-loading, teaching material preparation—FullFullFull
02Multimodal Interaction Architecture & Core ConceptsEdge vision, voice Agent, MCP protocol, and cloud-edge collaboration architecture analysis; technical differences and selection criteria between cloud-collaboration and local-offline deployment formsSenseCraft AIFullFullFull
03On-Device Lightweight Vision Inference ExperienceWatcher object detection model experience (material recognition, person approach, specific action perception), SenseCraft AI zero-code vision model adaptation, UART and network data output format analysisSenseCAP WatcherFullFullFull
04Scenario-Based Voice Q&A & Agent MechanismWarehouse/retail scenario natural-language real-time query demo, voice input→inference→TTS announcement full-pipeline experience, Agent prompt and role setting, dialogue memory mechanism comparison (no memory/short-term/long-term)SenseCraft AI AgentFullFullFull
05MCP Tool Calling & Business Integration DemoMCP-protocol-based external tool calling demo (real-time warehouse database query), smart warehouse full-pipeline demo (voice inventory query/stock-in/stock-out), OpenClaw desktop automation linkage demoMCP / OpenClawFullFullFull
06Watcher Vision AI & On-Device Network ConfigurationWatcher network configuration and SenseCraft platform binding, vision model selection and confidence/trigger threshold parameter tuning, event reporting rule configuration (object appearance, zone detection)SenseCraft AI / WatcherNoneFullFull
07Voice Agent & Role Prompt ConfigurationDefine Agent roles (warehouse assistant/exhibition guide), configure System Prompt and interaction style, memory mode switching and effect verificationSenseCraft AI AgentNoneFullFull
08Local Business Management System Docker DeploymentDeploy sample WMS (suharvest/warehouse_system) on RK3588-40 using Docker and Git, access management console to complete admin initialization and API Key generation, import demo materials and location dataDocker / WMSNoneFullFull
09MCP Bridge Service Configuration & Integration TestingObtain Watcher MCP access endpoint and authentication info, edit config.yml to configure local business system API address and API Key, start MCP Bridge service and verify endpoint handshake statusMCP BridgeNoneFullFull
10OpenClaw Automation Task LinkageDeploy OpenClaw automation tool environment, register automation scripts as MCP-callable tools, configure voice-command-triggered automated queries and scheduled tasksOpenClawNoneFullFull
11Full-Pipeline Integration Testing & TroubleshootingExecute typical business command integration testing (inventory query, stock-in submission, logistics tracking), troubleshoot common network timeout/API authentication failure/port conflict issues—NoneFullFull
12Local Offline Voice Pipeline Architecture AnalysisVAD/ASR/LLM/TTS module responsibilities and data flow latency analysis, offline vs cloud solution metric comparison (end-to-end latency, concurrency limits, VRAM usage, and privacy compliance)—NoneNoneFull
13Jetson Runtime Environment & Model Deployment TuningVerify JetPack/CUDA/TensorRT/PyTorch runtime environment, deploy ASR speech recognition model (Whisper/FunASR), deploy 4-bit quantized local LLM (Qwen2.5-7B-Instruct) and TTS engine (ChatTTS/Piper), route Watcher audio stream to local service portJetson Orin NXNoneNoneFull
14Offline Integration Testing & Latency OptimizationPhysically disconnect external network to verify LAN self-closed-loop operation, test each stage latency, tune model context length and sampling parameters—NoneNoneFull
15Solution Review and Delivery SummaryGroup result presentation and business scenario adaptation defense, cloud SaaS architecture vs local edge computing architecture cost and selection retrospective, business system API extension specifications and standardized delivery document archiving—PartialFullFull

● Full◐ Partial— None●+ Extended

The coverage key maps to course format IDs (taster / workshop / bootcamp), with values of full (complete coverage) / part (abbreviated coverage) / none (not included) / plus (deeper than full version). The taster session focuses on L1 on-device experience and MCP tool calling demo, excluding local business system deployment and offline pipelines; the workshop covers full L1+L2 Watcher configuration, WMS deployment, and MCP bridging; the bootcamp fully covers L1+L2+L3, including Jetson offline voice pipeline deployment.

Pick the layer, then the format

Time and goals determine which layer to choose.

  • Taster Session

    No FP

    1 day · 6–8h · L1 presentation layer · focusing on on-device experience and MCP tool calling demo

    • Day 1 MorningModules 01 + 02 + 03

      Environment pre-check → Multimodal architecture concepts → On-device vision inference experience

    • Day 1 AfternoonModules 04 + 05 + 15 (abbreviated)

      Voice Q&A and Agent mechanism → MCP tool calling and business integration demo → Summary review

    The taster session goal is "understand, explain, and demonstrate" — achieve the demo effect of voice query and vision-triggered linkage in 3 minutes. Does not include local business system deployment or offline voice pipelines.

  • Hands-On Course

    Full FP

    2–3 days · 14–20h · L1+L2 · Watcher configuration + local WMS deployment + MCP bridging + automation linkage

    • Day 1Modules 01–05

      Environment pre-check → Multimodal architecture → On-device vision → Voice Agent → MCP tool calling demo

    • Day 2Modules 06–09

      Watcher vision configuration → Agent role prompts → WMS Docker deployment → MCP bridge integration testing

    • Day 3 (optional)Modules 10 + 11 + 15

      OpenClaw automation linkage → Full-pipeline integration testing → Solution review and delivery summary

    The workshop delivers one complete multimodal system integration test including vision perception, voice Agent, local WMS, and automation tools. Student prerequisite: Docker fundamentals and REST API calling experience.

  • Delivery Course

    Full FP

    3–5 days · 24–35h · L1+L2+L3 · full coverage including Jetson offline voice pipeline deployment and offline verification

    • Day 1–2Modules 01–11

      Full L1+L2 content (on-device experience + Watcher configuration + WMS deployment + MCP bridging + full-pipeline integration testing)

    • Day 3Module 12

      Local Offline Voice Pipeline Architecture Analysis

    • Day 4Module 13

      Jetson Runtime Environment & ASR/LLM/TTS Model Deployment Tuning

    • Day 5Modules 14 + 15

      Offline integration testing and latency optimization → Solution review and delivery archiving

    The bootcamp goal is the ability to deliver AI interaction solutions in high-privacy and industrial isolated-network environments. Student prerequisite: Linux, PyTorch/Jetson fundamentals, and shell operation skills, familiarity with L1–L2 competencies.

The taster session is the standard format for solution demos and client communication: zero deployment barrier, 1-day closed loop, focusing on "voice can query, vision can trigger." Suitable for exhibitions, technology open days, and initial client contact scenarios.

Workshop Day 3 is an optional flexible day: if students have a strong foundation, it can be compressed to 2 days (Day 2 afternoon merged with OpenClaw and full-pipeline integration testing); if more MCP bridge tuning time is needed, use the full 3 days.

L1/L2 cloud collaboration solutions require stable uplink internet bandwidth on site for SenseCraft AI platform and LLM service calls; L1/L2 cannot run in external-network-free environments and must switch to the L3 offline solution.

Speech recognition accuracy is affected by on-site ambient noise, dialect accents, and specialized industry vocabulary; high-noise industrial sites require directional audio capture equipment, and far-field blind capture effects are not promised.

The L3 pure local solution relies on 100 TOPS-class edge GPU hardware (Jetson Orin NX 16GB and above), the 6 TOPS RK3588-40 cannot host local LLM inference.

The taster session does not include local business system deployment or offline voice pipelines. Do not promise clients that taster session students can independently complete WMS deployment or offline pipeline setup — that is the workshop and bootcamp delivery standard.

Not a pile of demos, but deliverables that can be lit up, validated, and replicated

The following are full-version (bootcamp) deliverables; the workshop delivers the first 4 items; the taster session delivers abbreviated versions of items 1 and 2.

  • 01 Watcher Hardware Configuration & Vision Model Parameter Specification

    Includes Watcher device network setup records, vision model selection and confidence/trigger threshold parameters, event reporting rule configuration checklist.

  • 02 Multimodal Agent role prompts and memory strategy configuration files

    Includes Agent role System Prompt, interaction style settings, dialogue memory mode (no memory/short-term/long-term) configuration and effect verification records.

  • 03 Local Business System & MCP Bridge Service Deployment Guide

    Includes Docker Compose deployment files, WMS management console initialization steps, MCP Bridge config.yml configuration template, and API Key management specifications.

  • 04 OpenClaw automation tool configuration scripts

    Includes OpenClaw environment deployment steps, configuration for registering automation scripts as MCP-callable tools, and configuration examples for voice-command-triggered automated queries and scheduled tasks.

  • 05 Local offline voice AI pipeline deployment and tuning manual (L3)

    Includes VAD→ASR→LLM→TTS module deployment steps, Jetson VRAM allocation and quantized model optimization parameters, offline integration test records, and end-to-end latency test report.

Who This Course Is For

Solution consultants and business sales Hands-on instructors and curriculum development teams Enterprise informatization and intelligence engineers Faculty and students at vocational colleges and applied universities

The value of this course is not in the LLM, but in the method of "integrating business systems into voice interaction"

M2 is not a course that teaches students to "chat with AI," but a methods course teaching teams how to use physical AI terminals and standard protocols to integrate already-existing WMS/ERP/CRM systems on site into natural-language interaction. What Chaihuo delivers is never just "one class session," but a complete set of things that can be taken apart, rewritten, and reassembled: 15-module course skeleton, teacher lesson plans and PPT, Watcher configuration templates, MCP bridge config.yml examples, Docker Compose deployment files, offline voice pipeline deployment manual.

  • Opening 01

    Change the Scenario

    The query content of Module 04 "Scenario-Based Voice Q&A" is open: your industry, your client site, a real problem happening in this city. Querying inventory can be warehouse management, exhibition exhibits, or meeting room schedules — the closer the problem is to a real site, the better the effect, and you know this better than we do.

  • Opening 02

    Connect Systems

    Your existing client business systems, management software on school training platforms, and partner REST API services can be connected after Module 08 to become the object pool for MCP bridging practice. M2 is responsible for explaining the method thoroughly; what systems to connect behind the door is up to you.

  • Opening 03

    Add Your Own

    What you have accumulated in the industry: Agent prompt tuning experience, MCP authentication pitfalls encountered, the analogy that makes students instantly understand voice pipeline latency, the three privacy questions most commonly asked at client sites — those are precisely the parts we do not have and cannot provide.

The best destiny of an AI interaction course is not to be executed in full once, but to be modified beyond recognition by an engineer and then become the solution that only he can deliver.
— Feng Lei, Author of This Course Series

Scope Boundaries & Compliance

Core Principles

L1/L2 business data flows within the LAN via local MCP bridging, core data stays in-domain; L3 runs purely local offline, zero public-network dependency.

In Scope

  • On-Site Target Perception & Structured Business Voice Q&A
  • MCP-Protocol-Based WMS/ERP/CRM System Tool Calling
  • Edge Vision Event Triggering & Automation Task Linkage
  • Multimodal Interaction for Smart Warehousing, Exhibition Guidance, Smart Reception, and Other Scenarios
  • Pure Local Offline Voice Interaction in High-Privacy & Industrial Isolated-Network Environments (L3)

Out of Scope

  • Does not replace high-concurrency, long-chain dedicated human customer service systems
  • Does not promise 100% accurate inference for ambiguous subjective multi-turn logic
  • Does not include reverse-engineering development for closed legacy systems without customer-open APIs
  • Not applicable to far-field blind capture in high-noise industrial sites (directional audio capture required)
  • L1/L2 solutions rely on internet connection to LLM services; external-network-free environments must switch to the L3 offline solution
  • L3 offline solution requires 100 TOPS-class edge compute (Jetson Orin NX 16GB and above), RK3588-40 (6 TOPS) cannot host local LLM inference
  • Speech recognition accuracy is affected by on-site ambient noise, dialect accents, and specialized industry vocabulary; specific recognition accuracy metrics for particular scenarios are not promised

Course Combinations Including This Module

M2 Multimodal AI Interaction

Partner With Us