What's New

At a launch event witnessed by more than ten embodied-AI hardware makers, PhanthyMotus opened its developer incentive programme: a shared hardware pool for remote debugging and free compute tokens for contributors.

Object recognition runs TensorRT directly on YOLOE-26s: 31.9 → 19.8 ms per frame on an Orin 5, accuracy from LVIS 24.4 to 30.8. A new visual_depth card estimates depth from an ordinary camera — in metres, not a relative scale — and a tape measure against a flat wall calibrates it to your own camera. Vision cards also answer about a single photo or URL with no camera attached.

LLM calls are streamed, so a stuck connection and a long answer are finally distinguishable. Liveness is judged on time-to-first-token (10/20/60 s, rising per retry) and the gap between tokens, cutting the worst case from six minutes to 90 seconds — while the total budget actually loosens to 600 s so long answers are no longer killed by mistake.

Results are now parsed by their own type. Image search actually returns image URLs (they used to be dropped entirely), video results carry duration, resolution and cover, and WebFetch can save to disk — which completes the search → download → send chain that used to break on the last metre.

Progress narration moves from "the model volunteers a sentence" to a framework guarantee: past a threshold of silent rounds or seconds, it generates a short spoken progress line. It skips when there is genuinely nothing new, and immersive skills can ask for quiet.

High-priority input now wakes the agent out of a wait on an unrelated running action — a "stop talking" said mid-narration used to sit for up to 58 seconds. Narration itself also queues instead of cutting in. Subagent round budgets rise from a hardcoded 10 to a configurable 50, and running out now writes up what it found instead of discarding it.

A RealSense D435 card publishes colour, depth and infrared from one capture pipeline shared across card instances, found by serial number rather than a renumberable device path. A two-finger gripper card takes 0–1000 over the arm’s existing SDK connection and refuses out-of-range values outright.

A new parakeet-en model trained on ~1.7M hours of diverse audio — deliberately including non-speech material to suppress hallucination — replaces an audiobook-trained option that was a poor match for a robot microphone. Punctuation and casing included, 104 MB, RTF 0.039 on CPU.

Japanese synthesis now streams instead of waiting for the whole utterance. On CPU a 26-second sentence starts after 14.4 seconds instead of 25; on GPU it runs at RTF 0.069. The move also removes a hazard that could kill the entire perception process on JetPack 5.11.

The PNDbotics Adam gains three read-only telemetry cards — per-joint motor state, whole-robot state, and 12-channel feedback per hand at 50 Hz — plus an emergency stop. Its communication interface is now configurable for wired sites.

The platform's first underwater robot arrives with six cards: attitude/depth/position, status, battery, IMU, a live RTSP camera and six-degree-of-freedom motion control. MAVLink over UDP, with a one-second heartbeat that refuses to move if the link has been quiet for three seconds.

Two new speech engines: kokoro-multi covers nine languages with a runtime dropdown, and mms-th adds Thai. Engines are renamed <model>-<languages>, with the old names still resolving. Numbers in English sentences are no longer read in Chinese, and re-applying a config no longer costs seconds per utterance.

A new face_recognition card enrolls a face from a photo or a link and then says who is in frame, with a confidence figure. Inference and the face database both stay on the robot; JetPack 5.11 machines run it on the GPU.

The AgiBot X2 joins the platform with 23 MCP cards — locomotion, 26 preset motions, dexterous hands, Linkcraft motion library, RGB-D + lidar + SLAM, TTS, expressions and LED.

Five robots in one building and none of them knows the others exist. Putting them on one ROS domain is the obvious move and also the dangerous one: a command typed on one robot was executed by a second, same timestamp in both logs. This is the link we built instead — key-based identity, Bluetooth-style six-digit pairing, signed requests against a pinned key, and three gates before a peer can move anything. Including the scars: a keyword-based permission check that let a read-only peer drive locomotion, and a config flag that looks like a safety gate but is read by no code.

Anthropic opened a research preview of the Model Hardware Standard. Starting from a microscope rig and from a walking humanoid, two teams converged on four judgments and diverged on three constraints. A grounded comparison of the problems, the logic, and what is actually shipping.

The RealMan RM75-6F-V driver is available. Seven joints in one card, a pre-move self-check that refuses unsafe commands instead of moving first, motion that only reports done when the arm has actually arrived, a full set of live state cards, and a 3D skeleton view.

The EngineAI T800 driver grows a speaker, walking with real gaits, a motion recorder, and a posture switch that always routes through standing — plus honest state reporting instead of guesses.

A G1 that squatted could not stand back up: the stand-up call reported "robot is on ground" and bounced between states. Fixed on the real robot, along with the deeper confusion between control mode and physical posture.

The Booster K1 bipedal humanoid now has a first-class driver: walking, gait and mode switching, head aiming, upper-body joint control, preset actions, get-up, trajectory playback, audio and vision — all as tools your agent can call.

The BrainCo Revo 2 dexterous hand now has a driver: open and close, per-finger positions, preset gestures, live joint telemetry, LED and button events, and fingertip touch on the tactile model — with two hands on one bus.

Feishu channels now work in group chats as well as direct messages, replies land on the message that triggered them, images and files come through, other bots can be allowed in one at a time, and roles are managed from a panel in the dashboard.

The RobotEra Q5 driver adds waist rotation as a first-class control, and fixes the arm gestures: no more mirrored high-five, no more jerky motion, and a reset that brings both arms home together.

The ROBOTERA Q5 wheeled humanoid is now supported: dual-arm manipulation with 11-DOF dexterous hands, a liftable torso, and LiDAR-based autonomous navigation.

The PNDbotics Adam humanoid now has a stereo depth camera and per-finger hand control — it can judge how far away something is and shape its hand to pick it up.

The T800 can now turn and tilt its head to face whoever is speaking, and greet people with a wave or a handshake — the small gestures that make an interaction feel human.

Wiring up ASR, an LLM and TTS takes an afternoon. Holding a conversation that gets interrupted is a different problem. Three failures from the field — you cannot get a word in, a grunt becomes a command, the answer finishes before the action does — and the mechanisms each one forced: priority routing, three interrupt modes, and a completion protocol. Including one knob we found is not actually wired up.

Speech recognition can now run on the GPU for roughly 2.5× faster turnaround per utterance. It's opt-in, because the speed costs memory — the default stays on CPU.

The canvas now shows when someone else is editing, so two people working on the same robot stop silently undoing each other's changes.

The robot's voice sounds noticeably more human, and you can now choose which speech engine it uses to trade off naturalness against speed and memory.

Tuning a robot takes hours. Now you can save the whole working setup as a solution, load it onto another robot in one click, or pick a ready-made one from the marketplace.

Point the camera at a sign, a label, or a door plate and the robot reads it. Recognition runs on the robot itself, so no image ever leaves the machine.

A single account entry ties the dashboard to your skills, solutions, and community identity — and it works on a phone, so you can install a skill without sitting at the robot.

Conversations are no longer text-only. Send a photo of a shelf and ask what's wrong with it, or get the robot to send back what it just saw — images, video, and documents now travel both ways.

The industrial-grade Lynx M20 quadruped joins the platform. Four gaits, dual lidar, dual cameras, and autonomous charging — all controllable through natural language.

The robot's AI now processes half the context per request and hits cache twice as often. Responses come faster, costs go down, and background monitoring gets smarter.

Depth cameras stream 47x smaller, speaker audio no longer stutters, the AI responds faster, and the robot won't repeat itself after speaking.

The 25-43 DOF PNDbotics Adam humanoid joins the platform. Walk, manipulate objects with dexterous hands, and visualize full-body motion in 3D — all through natural language.

The industrial-grade Matrice 300 RTK drone joins the platform. Fly waypoints, stream 6-direction perception, and monitor aircraft health — all through natural language.

The 21-DOF Noetix Bumi-EDU mobile manipulator joins the platform. Walk, navigate, talk, and perceive depth — a compact humanoid for education and research.

Built for embodied scenarios like guided tours and long-term companionship — the AI now compresses context intelligently so it never loses track, even after hundreds of exchanges.

The 25-DOF EngineAI T800 humanoid joins the platform. Walk, dance, gesture, and speak — all controllable through natural language.

Some actions no longer wait for the AI to think. LED changes, speech interrupts, and status queries now respond in under 50ms.

The AI plans ahead while the robot is still moving. Actions execute asynchronously for speed, but logical order is always preserved — no more conflicts or confusion.

Browse community skills, install with one click, or publish your own. No coding needed — skills are defined in natural language.

Send a message on Feishu, your robot reads it, thinks, and replies — all through your existing team chat.

The dashboard is now fully usable on mobile — native tab bar, touch-friendly canvas, and all features accessible from your pocket.
Update Agent Core, Perception, and drivers from the dashboard. Choose between preview, release, and stable channels.
Fly a drone manually once, record the path, and replay it autonomously as many times as you want.

The classic Unitree Go1 quadruped joins the supported hardware list — same AI capabilities, one more robot.

Full AI agent control for DJI Mavic 3E/3T and 4E/4T — fly, record waypoints, switch cameras, and measure temperature by voice.
Speech recognition, voice synthesis, and wake-word detection all run locally. Your robot works the same in a basement or a field.
Real-time 3D visualization in the browser: LiDAR point clouds, SLAM maps, humanoid skeleton, and quadruped pose — all live.

No more SSH to change networks. Scan, connect, and manage WiFi right from the web interface on your robot.

Dashboard now runs over HTTPS with auto-generated certificates. Access your robot securely from any browser, anywhere.
Use your phone or laptop's microphone to talk to the robot directly through the dashboard — no extra hardware needed.
Tell your robot what to look for in natural language. YOLO + CLIP runs on Jetson GPU — no cloud, no pre-defined classes.
Your robot now automatically slows down and stops before hitting obstacles. Works with G1, Go2, and any robot with LiDAR.

Deploy a complete AI agent on your robot in minutes. No vendor lock-in, no cloud dependency — everything runs on-device.