跳到正文

UIseries

ai-robot

RK3506 Voice Robot An embedded AI voice assistant running on the Rockchip RK3506 development board. Pure C, single-threaded event loop, zero dynamic memory allocation.

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

<<<<<<< HEAD

🎙️ RK3506 Voice Robot

An embedded AI voice assistant running on Rockchip RK3506 — push a button,ask a question, and see the AI’s answer rendered on an LCD screen.

Button Press → Voice Recording → WiFi Upload → AI Agent → LCD Display


图片:robot

Table of Contents

  • Overview
  • System Architecture
  • State Flow
  • Hardware Requirements
  • Quick Start
  • Build Guide
  • Configuration
  • AI Server API
  • Project Structure
  • Tech Stack
  • Design Decisions
  • License

Overview

RK3506 Voice Robot is a pure-C embedded application that turns a Rockchip RK3506 development board into a voice-driven AI companion. It captures speech via a microphone, uploads the audio over WiFi to a remote AI Agent, and renders the text reply on an LCD screen — all without any dynamic memory allocation.

Why this project?

Use CaseDescription
🧒 Kids’ Q&APress a button, ask “How did dinosaurs go extinct?”, see the answer
🍳 Kitchen HelperQuery recipes with messy hands — “How do I make tomato egg stir-fry?”
👴 Elderly CompanionVoice-first interface for seniors uncomfortable with smartphones
💼 Desk AssistantQuick queries during work — exchange rates, weather, reminders

Key Metrics

MetricTarget
Voice upload → screen response≤ 3 seconds (excluding recording time)
Memory allocation100% static (compile-time), no malloc/free at runtime
WiFi resilienceAutomatic reconnection on disconnect
Error handling5 categories: audio / WiFi / HTTP / timeout / server error
Single-utterance recordingUp to 60 seconds at 16kHz/16bit/mono

System Architecture

┌──────────────────────────────────────────────────────────────────┐
│                      main.c  (50ms event loop)                    │
│                                                                    │
│  ┌───────────┐    ┌──────────────────┐    ┌───────────────────┐  │
│  │ ButtonLED │◄───│   StateMachine    │───►│   AudioCapture    │  │
│  │ (GPIO)    │    │                   │    │    (ALSA PCM)     │  │
│  └───────────┘    │  IDLE             │    └───────────────────┘  │
│                   │  RECORDING        │                            │
│                   │  PROCESSING       │    ┌───────────────────┐  │
│                   │  RESULT / ERROR   │───►│   HttpClient      │  │
│                   └────────┬─────────┘    │   (libcurl)       │  │
│                            │              └─────────┬─────────┘  │
│                            ▼                        │            │
│                   ┌──────────────────┐    ┌─────────┴─────────┐  │
│                   │    Display       │    │   WiFiManager     │  │
│                   │ (Framebuffer +   │    │ (wpa_supplicant)  │  │
│                   │  FreeType)       │    └───────────────────┘  │
│                   └──────────────────┘                            │
└──────────────────────────────────────────────────────────────────┘

Layered Design

LayerModuleResponsibility
🎮 Applicationmain.cInit all modules → 50ms poll loop → graceful shutdown
🧠 Business Logicstate_machine9 transition rules, 5 states with enter/update/exit handlers
🔌 Core Servicesaudio_capture, http_client, display, wifi_managerIndependently testable modules
🖥️ Hardware DriversALSA, Framebuffer, GPIO sysfs, SPILinux kernel interfaces

State Flow

                    ┌─────────────────────────────────────┐
                    │                                     │
                    ▼                                     │
              ┌──────────┐  btn_press   ┌──────────────┐ │
              │   IDLE   │ ──────────► │  RECORDING   │ │
              │          │ ◄────────── │              │ │
              └────┬─────┘   timeout   └──────┬───────┘ │
                   ▲                          │          │
                   │              btn_release │          │
                   │              or timeout   │          │
                   │                          ▼          │
                   │                    ┌──────────────┐ │
                   │                    │  PROCESSING  │ │
                   │                    └──┬───────┬───┘ │
                   │               success│       │fail  │
                   │                      ▼       ▼      │
                   │          ┌────────────┐ ┌─────────┐ │
                   └──────────│   RESULT   │ │  ERROR  │ │
                    timeout / │            │ │         │ │
                    btn_press └────────────┘ └─────────┘ │
                                                         │
                              └──────────────────────────┘
StateScreenLEDBehavior
IDLE”Press button to ask”OFFPolls GPIO button
RECORDING”Listening…” + durationONALSA capture to static buffer (max 60s)
PROCESSING”Thinking…”FAST BLINKWiFi check → HTTP upload → parse JSON reply
RESULTAI reply textOFFDisplay answer for 5s, then back to IDLE
ERRORError icon + descriptionOFFDisplay error for 3s, then back to IDLE

Hardware Requirements

ComponentSpecificationNotes
SoCRockchip RK3506ARM Cortex-A53
OSLinux (Buildroot / Yocto)Kernel 5.10+, glibc or musl
MicrophoneI2S or USB micALSA hw:0,0 compatible
LCDSPI / RGB interfaceDefault 240×320, Framebuffer /dev/fb0
WiFiSDIO / USB modulewpa_supplicant managed
ButtonGPIO push buttonPull-up, active-low or active-high
LEDGPIO LEDStatus indicator

💡 The GPIO pin numbers for the button and LED are configured in config/app_config.ini. Defaults are placeholders — adjust to match your board.


Quick Start

1. Clone & prepare dependencies

git clone https://github.com/yourname/rk3506-voice-robot.git
cd rk3506-voice-robot

On your cross-compilation host, install build dependencies:

# Debian/Ubuntu
sudo apt install libasound2-dev libcurl4-openssl-dev \
                 libfreetype6-dev libconfig-dev cmake

# Also install the aarch64 cross-compiler
sudo apt install gcc-aarch64-linux-gnu

2. Download a Chinese font

Place a TrueType Chinese font (e.g., Noto Sans SC Regular) in resources/:

# Example — rename your downloaded font
mv ~/Downloads/NotoSansSC-Regular.ttf resources/

3. Configure

Edit config/app_config.ini — at minimum, set your WiFi credentials:

[wifi]
ssid = "YourWiFiSSID"
psk  = "YourWiFiPassword"

4. Cross-compile

mkdir build && cd build
cmake .. -DTOOLCHAIN_PATH=/usr -DCMAKE_BUILD_TYPE=Release
make -j$(nproc)

The output is build/voice_robot — a statically-linkable ARM aarch64 binary.

5. Deploy & run

Copy the binary, config, fonts, and scripts to your RK3506 board:

# On the target device
./scripts/start_wifi.sh     # Bring up WiFi
./scripts/run_voice_robot.sh

Build Guide

Cross-compilation (aarch64 target)

mkdir build && cd build

# Option A: system cross-compiler
cmake .. -DCMAKE_BUILD_TYPE=Release

# Option B: custom SDK toolchain
cmake .. -DTOOLCHAIN_PATH=/opt/rk3506-sdk \
         -DCMAKE_BUILD_TYPE=Release

make -j$(nproc)

CMake options:

OptionDefaultDescription
TOOLCHAIN_PATH(empty)Path to aarch64 toolchain root (bin/aarch64-linux-gnu-gcc)
CMAKE_BUILD_TYPE(empty)Release or Debug

Local build (x86_64 for testing)

mkdir build && cd build
cmake .. -DCMAKE_BUILD_TYPE=Debug
make -j$(nproc)

⚠️ Local builds are useful for syntax checking and static analysis, but GPIO, ALSA, and Framebuffer calls will fail without real hardware.

Build artifacts

build/
├── voice_robot        # Main executable
├── CMakeFiles/
└── CMakeCache.txt

Configuration

All runtime parameters live in config/app_config.ini. The application reads this file once at startup.

Full reference

SectionKeyTypeDefaultDescription
[audio]devicestringhw:0,0ALSA PCM capture device
sample_rateint16000Sampling rate in Hz
channelsint11 = mono, 2 = stereo
record_duration_secint60Maximum recording duration
[wifi]interfacestringwlan0Wireless interface name
ssidstringyour_ssidWiFi network SSID
pskstringyour_passwordWiFi password (WPA2-PSK)
[http]server_urlstringhttp://47.81.10.119/AI Agent server URL
timeout_secint10HTTP request timeout in seconds
[display]fb_devicestring/dev/fb0Framebuffer device path
font_pathstringresources/NotoSansSC-Regular.ttfTrueType font path
font_size_ptint18Font size in points
lcd_widthint240Screen width in pixels
lcd_heightint320Screen height in pixels
[gpio]button_gpioint0Button GPIO number (sysfs)
led_gpioint1LED GPIO number (sysfs)
[system]log_levelint20=ERROR, 1=WARN, 2=INFO, 3=DEBUG

🔄 Changes take effect after restarting voice_robot.


AI Server API

The AI Agent server is a separate service that receives audio and returns text replies.

Audio Upload

POST {server_url}
Content-Type: multipart/form-data
FieldTypeDescription
audiofileWAV file, 16kHz / 16-bit / mono PCM

Response Format (Success)

{
    "reply": "The AI assistant's response text"
}
  • HTTP 200–399: success — the reply field is extracted and displayed on screen.
  • HTTP 4xx/5xx: failure — an error message is shown with the HTTP status code.

Testing with curl

# Test the AI endpoint directly
curl -X POST http://47.81.10.119/ \
     -F "audio=@test.wav"

Project Structure

rk3506_voice_robot/
│
├── CMakeLists.txt                  # Build system (cross-compile aarch64)
├── README.md                       # This file
│
├── config/
│   └── app_config.ini              # Runtime configuration (6 sections, 20 fields)
│
├── include/
│   ├── common_types.h              # Shared enums, structs, callbacks, LOG macros
│   ├── error_codes.h               # Error codes (-1xxx to -9xxx by module)
│   ├── app_config.h                # Config parser API (load/reload/error)
│   ├── audio_capture.h             # ALSA PCM capture (open/set_params/read/close)
│   ├── wifi_manager.h              # WiFi status detection + reconnect
│   ├── http_client.h               # HTTP multipart upload + JSON parse callback
│   ├── display.h                   # Framebuffer drawing + text rendering
│   ├── font_renderer.h             # FreeType glyph rendering + UTF-8 decode
│   ├── state_machine.h             # 9 transition rules, 5-state controller
│   └── button_led.h                # GPIO sysfs button polling + LED control
│
├── src/
│   ├── main.c                      # Entry point: init → 50ms event loop → shutdown
│   ├── app_config.c                # libconfig parser with validation
│   ├── audio_capture.c             # ALSA capture (xrun recovery included)
│   ├── wifi_manager.c              # /sys/class/net/operstate reader (zero-fork)
│   ├── http_client.c               # WAV header assembly + multipart POST + cJSON parse
│   ├── display.c                   # Framebuffer pixel ops (32bpp + 16bpp/RGB565)
│   ├── font_renderer.c             # FreeType init → FT_Load_Glyph → GlyphBitmap
│   ├── state_machine.c             # Full enter/update/exit for all 5 states
│   └── button_led.c                # sysfs GPIO export/direction/value I/O
│
├── libs/
│   └── cJSON/
│       ├── cJSON.h                 # v1.7.18 (ARM GCC compatible)
│       └── cJSON.c                 # ~2750 lines, single-file JSON parser
│
├── resources/
│   └── font_placeholder.txt        # Instructions for downloading Noto Sans SC
│
└── scripts/
    ├── start_wifi.sh               # wpa_supplicant init + dhclient
    └── run_voice_robot.sh          # LD_LIBRARY_PATH + exec voice_robot

Stats: 12 headers, 9 source files, 2 library files, 2 scripts, 1 config, 1 font placeholder = ~4,800 lines of C.


Tech Stack

ComponentTechnologyWhy
LanguageC11Zero runtime overhead, static memory, ideal for embedded
BuildCMake 3.16+Cross-compilation for aarch64 toolchains
AudioALSA libasoundLinux standard, RK3506 BSP includes driver
HTTPlibcurl (easy)Mature multipart/form-data support
JSONcJSON v1.7.18Single-file C library, no dependencies
FontFreeType 2Industry-standard glyph rendering
DisplayLinux FramebufferUniversal, no GUI toolkit needed
ConfiglibconfigLightweight INI-style parser
GPIOsysfsStandard Linux GPIO interface (no kernel module needed)
WiFiwpa_supplicantPre-installed on Buildroot/Yocto

Design Decisions

Why pure C and static memory?

RK3506 has limited RAM (~256MB–512MB). Dynamic allocation (malloc/free) leads to fragmentation over long-running sessions. All buffers are compile-time sized: the recording buffer is 16kHz × 2bytes × 60s = 1.92MB, and the HTTP response buffer is 4KB — both statically allocated.

Why a single-threaded event loop?

Multi-threading on embedded Linux introduces mutex/condition-variable complexity and context-switch overhead. The 50ms poll loop is simple, predictable, and sufficient for this I/O profile (button polling, audio drain, network I/O).

Why Framebuffer instead of a GUI toolkit?

The display only renders simple text (no widgets, no animations). Direct framebuffer access skips the heavy dependencies of toolkits like Qt or LVGL while supporting both 32bpp (ARGB) and 16bpp (RGB565) pixel formats — covering virtually all SPI LCDs on the market.

Error handling philosophy

Every module returns 0 on success and negative ErrorCode on failure. The state machine treats failures as events (EVENT_HTTP_FAIL) and routes them to a dedicated ERROR state with a user-friendly screen message — the robot never gets stuck.


Roadmap

  • P0 — Core voice → AI → display pipeline
  • P1 — Button/LED interaction, state machine, error handling
  • P2.1 — VAD (Voice Activity Detection) for hands-free mode
  • P2.2 — Conversation history with scrollback on screen
  • P2.3 — TTS (Text-to-Speech) audio playback
  • P2.4 — OTA firmware upgrade over WiFi

License

MIT © 2026


Built with ❤️ for the embedded Linux community

=======

ai-robot

RK3506 Voice Robot An embedded AI voice assistant running on the Rockchip RK3506 development board. Pure C, single-threaded event loop, zero dynamic memory allocation.

33334bbbfbdf258c17264f132e553142d668b8d3

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。