Arduino VENTUNO Q board controlling an SO-101 robot arm picking up a rubber duck with on-board AIArduino VENTUNO Q board controlling an SO-101 robot arm picking up a rubber duck with on-board AI

A $299 Arduino board just turned a desktop SO-101 robot arm into a machine that sees, decides and picks up rubber ducks with no cloud in the loop. The Arduino VENTUNO Q pairs a Qualcomm Dragonwing IQ8 AI processor with an STM32H5 real-time microcontroller, and it’s the clearest sign yet that physical AI has reached the embedded bench.

What the Arduino VENTUNO Q actually is

Arduino opened pre-orders on 25 August 2026, describing VENTUNO Q as its “most advanced platform to date” in its launch announcement. The core idea is a “dual-brain” architecture. One brain thinks and the other acts:

  • AI brain: a Qualcomm Dragonwing IQ8 (QCS8275) with an octa-core Kryo Gen 6 CPU, Adreno 623 GPU and a Hexagon Tensor NPU rated at up to 40 dense TOPS. It ships with Ubuntu preloaded.
  • Action brain: an STM32H5F5 (Arm Cortex-M33 at 250 MHz, 4 MB flash, 1.5 MB RAM) running the Arduino Core on Zephyr RTOS for deterministic control of motors, CAN-FD, PWM and GPIO.
  • Memory and storage: 16 GB LPDDR5 and 64 GB eMMC, plus an M.2 slot for NVMe Gen.4.
  • I/O for robots: three MIPI CSI camera connectors, 2.5 Gb Ethernet, Wi-Fi 6, Bluetooth 5.3, a CAN-FD PHY on a screw terminal, and built-in ROS 2 compatibility.

According to the official Arduino store page (checked 8 October 2026), VENTUNO Q is listed at an introductory price of $299 as a pre-order, with delivery quoted at three to four weeks.

How the two brains split the work

The Linux side runs the heavy models: LLMs, vision-language models, speech recognition and object detection. The MCU side handles anything with a hard deadline. Arduino’s product page says the two talk over an RPC bridge, so a Python app on Ubuntu can call functions in a sketch running on the STM32 without you wiring a second board over UART.

Diagram of the Arduino VENTUNO Q dual-brain architecture: Qualcomm Dragonwing IQ8 running AI models linked by an RPC bridge to an STM32H5 MCU driving motors

This is the same split that experienced robotics engineers already build by hand: a Linux computer for perception and planning, a microcontroller for PWM, encoder reads and safety interlocks. If you’ve ever weighed whether you need ROS or bare-metal code, VENTUNO Q’s answer is “both, on one PCB”. The practical win is less integration work. The discipline is unchanged: never let a Linux scheduler hiccup stall a motor command.

What runs out of the box

Arduino says its App Lab ships with NPU-optimised models via Qualcomm AI Hub, including Qwen 3 4B (LLM), Qwen 2.5 7B and Qwen 3 4B VLMs, Gemma 4 E2B and E4B, Whisper ASR, Melo and Piper TTS, YoloX small object detection and MediaPipe gesture recognition. You can also load GGUF models from Hugging Face through a llama.cpp Brick, or train your own in the integrated Edge Impulse Studio. App Lab is optional; VS Code, Docker and native Linux tools work too.

The viral demo: SO-101 plus SmolVLA, fully local

The demo getting attention comes from Dmitry Maslov of Hardware.ai. As Hackster reported, he used a VENTUNO Q as the brain of an SO-101 arm, a low-cost design with a servo at each joint. The arm picks and places rubber ducks using Hugging Face’s SmolVLA model, running entirely on the board’s Dragonwing IQ8.

The inputs are modest: images from two cameras (one overhead, one on the gripper) plus joint-position data. According to Hackster, Maslov taught the task with just 50 demonstrations, and performance was good “even without any real optimization”. Dev Gadgets’ write-up describes the joint commands as being handed to the STM32H5 for low-level motor control. That fits the board’s design, but the published coverage doesn’t detail exactly how the inter-processor link was implemented, so treat it as the intended pattern rather than a documented recipe.

How vision-language-action models work

A vision-language-action (VLA) model takes camera frames, a natural-language instruction and the robot’s current state, and outputs robot actions directly. There are no hand-written inverse-kinematics routines for each task. Hugging Face’s SmolVLA announcement explains the architecture:

  1. Perception and language: a SmolVLM2 backbone (a SigLIP vision encoder plus a SmolLM2 decoder) turns images and the instruction into tokens. Each frame is compressed to 64 visual tokens.
  2. Robot state: joint positions are projected into a single token and processed alongside the image and text tokens.
  3. Action expert: a compact transformer of about 100M parameters, trained with flow matching, generates an “action chunk”, a short sequence of future joint commands, without slow token-by-token decoding.

The whole model is about 450M parameters, pretrained on community LeRobot datasets that focus on SO-100-class arms. That is why a small fine-tuning set can work: the model already knows what a gripper and a table look like. With asynchronous inference, the robot executes the current chunk while the next is computed; Hugging Face reports about 30% faster task completion.

Arduino VENTUNO Q vs Arduino UNO Q

VENTUNO Q is the big sibling of the UNO Q, which uses the same dual-brain idea at a much lower price. MakeUseOf’s analysis calls the concept a Raspberry Pi and an ESP32 on one board. It also notes that UNO Q prices rose by roughly a third, which Arduino attributed to memory costs.

Feature Arduino UNO Q (2GB) Arduino VENTUNO Q
Application processor Qualcomm Dragonwing QRB2210, 4× Cortex-A53 @ 2.0 GHz Qualcomm Dragonwing IQ8 (QCS8275), octa-core Kryo Gen 6
AI acceleration Adreno GPU, DSP, 2× ISP Hexagon NPU, up to 40 dense TOPS
Real-time MCU STM32U585, Cortex-M33 up to 160 MHz STM32H5F5, Cortex-M33 @ 250 MHz
RAM / storage 2 GB LPDDR4 / 16 GB eMMC 16 GB LPDDR5 / 64 GB eMMC + M.2 NVMe
Cameras USB (UVC) webcam; MIPI via headers 3× MIPI CSI + USB cameras
Networking Wi-Fi 5, Bluetooth 5.1 Wi-Fi 6, Bluetooth 5.3, 2.5 Gb Ethernet
Form factor UNO (68.85 × 53.34 mm) 160 × 100 mm
US store price (8 Oct 2026) $59 $299 (introductory, pre-order)

Rule of thumb: UNO Q for lightweight vision and small models in the shield ecosystem; VENTUNO Q for multi-camera robotics, local LLMs and VLAs, where NPU throughput and memory decide what you can run.

What this means for embedded engineers

Physical AI doesn’t replace embedded skills; it raises the stakes on them. The arm is just servos, and reliability comes from the system around the model. A realistic path to try this:

  1. Start with the LeRobot stack. Install LeRobot with the smolvla extra and fine-tune lerobot/smolvla_base on your own recorded episodes. Hugging Face’s example uses --policy.path=lerobot/smolvla_base with a batch size of 64.
  2. Collect clean demonstrations. Keep camera placement fixed (overhead plus wrist, like the demo), vary object positions, and record consistent task instructions. Data quality matters more than episode count.
  3. Keep the control loop on the MCU. Let the Linux side send target joint positions at the policy rate, and let the STM32 interpolate, enforce joint limits, and run a watchdog that stops the arm if commands stop arriving.
  4. Tune the low-level loop first. A VLA can’t fix an oscillating servo. Our guide on designing a robot control loop without oscillation covers the fundamentals.
  5. Measure latency end to end. Timestamp camera capture, inference and MCU actuation to see where the time goes.

If your background is microcontroller ML, this is the natural next step beyond edge AI on microcontrollers. The skills that transfer are quantisation, memory budgeting, and deterministic I/O.

Key takeaways

  • The Arduino VENTUNO Q combines a 40-TOPS Dragonwing IQ8 running Linux with an STM32H5 real-time MCU on one board.
  • Hugging Face’s ~450M-parameter SmolVLA ran locally on it to control an SO-101 arm after 50 demonstrations.
  • VLAs map camera images, instructions and joint states straight to action chunks. You still own safety, timing and motor control.
  • UNO Q ($59, 2GB) suits lighter edge AI; VENTUNO Q ($299 introductory) targets multi-camera robotics and local LLM/VLA work.

FAQ

What is the Arduino VENTUNO Q?

It is Arduino’s dual-brain edge-AI board. A Qualcomm Dragonwing IQ8 processor with an NPU rated up to 40 dense TOPS runs Ubuntu and AI models, while an STM32H5F5 microcontroller runs Arduino sketches on Zephyr RTOS for real-time control.

How much does the Arduino VENTUNO Q cost?

As of 8 October 2026, Arduino’s US store lists it at an introductory price of $299 as a pre-order, with delivery quoted at three to four weeks.

Can the Arduino UNO Q run SmolVLA?

The published demo used the VENTUNO Q, not the UNO Q. Hugging Face says SmolVLA is small enough to run on a CPU, but the UNO Q’s 2–4 GB of RAM and lack of a dedicated NPU make the VENTUNO Q the realistic choice for responsive control.

What is a vision-language-action model?

A VLA is a neural network that takes camera images, a text instruction and the robot’s state as input, and outputs robot actions directly. SmolVLA does this with a compact vision-language backbone plus a flow-matching “action expert” that predicts short sequences of joint commands.

Want to build the embedded, robotics and AI skills behind projects like this? Explore the Educational Engineering Team courses at https://eduengteam.com/wp-content/uploads/2026/10/meta-muse-gadgets-esp32-ai-agent-diagram.jpg.

Leave a Reply