
On July 17, the 9th World Artificial Intelligence Conference (WAIC) officially opened at the Shanghai World Expo Center. As an annual flagship event for the AI industry, this year’s conference gathered large model developers, AI application vendors and hardware manufacturers from home and abroad. Leitech’s AI-focused new media arm Leitech AGI (leikejiagi) also dispatched a reporting team to Shanghai for on-site coverage.
Agent technology and embodied intelligence could be seen everywhere across this AI-focused exhibition. However, one booth stood out: it showcased no humanoid robots or conceptual AI products, only several devices resembling routers and mini PCs. Surprisingly, this exhibitor, MoreMoto Intelligence, featured logos of major conglomerates including Lenovo, Great Wall Motors and China Mobile on its backdrop wall.

(Photograph credit: Houmo Intelligence)
Simply put, while most industry players are racing to build smarter Agents, MoreMoto focuses on developing the underlying chips powering these systems.
A widely overlooked detail is this: although large models keep growing more capable, with parameters jumping from 7 billion (7B) to 70 billion (70B) and then to 700 billion (700B), the AI running on your smartphone operates on a fundamentally different architecture from that deployed in corporate data centers. Cloud-based large models deliver strong performance, yet every request requires transmitting data to remote servers thousands of miles away. This process brings a host of drawbacks: latency, high power consumption, Token usage fees, and privacy vulnerabilities.
That explains why the industry has dubbed 2026 the "Year Zero of On-Device AI". The market has seen an influx of AI PCs, AI smartphones, intelligent robots and other AI endpoints, yet all share a common pain point: how to pack sufficient computing power for large model operation into palm-sized hardware without overheating or draining batteries rapidly.

(Photograph credit: LeiTech on-site production)
This is exactly the problem MoreMoto Intelligence has been tackling. Its newly launched M50 chip features modest headline specs on paper: 10W power consumption, 160 TOPS of computing performance, and native support for large models ranging from 30B to 120B parameters. Readers may notice its 10W power draw is far lower than competing chips with equivalent TOPS ratings.
This efficiency stems from the M50’s revolutionary Compute-in-Memory (CIM) architecture. Traditional chips separate computing and memory units; data must be shuttled back and forth for every calculation. This constant data transfer often consumes more energy than the computation itself — an industry bottleneck known as the "memory wall". MoreMoto’s CIM design integrates storage and computation onto a single silicon die. Eliminating repetitive data movement drastically cuts power usage.

(Photograph credit: LeiTech on-site production)
Products built around this chip deliver impressive functionality. Take the Lenovo AI Host P7 as an example: palm-sized and weighing only 300 grams, the entire device consumes merely 30W, enough to run off an ordinary power bank. It can run 122-billion-parameter large models offline with an inference speed of 50 tokens per second. In plain terms, even without internet access, users can chat with the AI, draft documents and analyze spreadsheets with nearly identical performance to an online connection. Most critically, privacy risks are eliminated, since all data processing occurs locally within the hardware with no online transmission leaks.

(Photograph credit: LeiTech on-site production)
The Great Wall N90 Pro is another standout device: a fully domestically produced AI laptop combining a Phytium CPU, Kylin OS, M50 chip and homegrown large models. It runs 35B-parameter models offline at an inference speed of 30 tokens per second, making its value self-evident for government, finance, energy and other regulated sectors. Meanwhile, Ququ Intelligent’s ClawHouse X1 Pro is a holographic interactive personal computing hub capable of simultaneous 3D rendering, inference for hundred-billion-parameter models and low-latency voice interaction, packaging smart office and smart education functions into a compact "personal AI server". The Lenovo AI Workmate concept device adopts a desktop robot form factor; it scans documents to auto-generate PowerPoint slides and projects content onto shared displays, with all features fully functional offline.

(Photograph credit: LeiTech on-site production)
Products showcased at WAIC highlight that MoreMoto pursues a vastly different roadmap from industry giant NVIDIA. NVIDIA builds general-purpose computing platforms designed for universal compatibility across all scenarios. MoreMoto develops chips purpose-built for on-device inference, enabling plug-and-play integration for hardware OEMs. Device manufacturers no longer need to develop chips from scratch or tolerate cloud API latency and recurring costs; AI upgrade is completed simply by installing an accelerator card. This collaborative model will accelerate the mass adoption of on-device AI far faster than many anticipate.
Why does on-device AI matter? Many believe its only benefit is offline operability, but this is merely the most superficial advantage. Its true value falls into three tiers:
The first is cost. Cloud-based LLM token fees add up quickly: average users spend ¥100–¥200 monthly, while heavy users easily exceed ¥400–¥500. Local inference eliminates recurring token charges—and even electricity costs remain negligible, slashing total cost of ownership. The second is privacy: your chat logs, work files, and personal data never leave your device—processing occurs entirely locally. You no longer need to trust any company’s “privacy pledge”; data never physically departs your device. The third is user experience: cloud AI inevitably incurs tens to hundreds of milliseconds of latency, compounded by routine network fluctuations. On-device inference achieves millisecond-level response times, enabling silky-smooth interactions—and your AI assistant isn’t competing with millions of concurrent users for compute resources. It’s truly yours: always available, instantly responsive.
The M50 currently boasts China’s most complete closed-loop ecosystem for domestic on-device AI chips, covering independent R&D, mass manufacturing and ecosystem development. Critically, its ability to hit 160 TOPS at 10W power cannot be achieved by cramming in extra transistors; the native CIM architecture inherently optimizes performance for edge deployments. Even more notably, MoreMoto sells more than just chips. It provides M.2 accelerator cards, add-on acceleration hardware and full software toolchains. Customers avoid tedious parameter tuning and model quantization, achieving plug-and-play deployment. This out-of-the-box usability significantly lowers barriers to large-scale commercialization.

(Photograph credit: LeiTech on-site production)
Nevertheless, the on-device AI market is still in its infancy. Though 2026 marks its official Year Zero, long-term growth hinges on two factors. First, mainstream hardware brands including Lenovo and Great Wall must achieve genuine market traction with their M50-powered products — advanced hardware means nothing if sales fail to materialize. Second, cloud AI costs continue to decline; sustained product iteration is required to preserve on-device AI’s competitive advantages long-term.
MoreMoto also faces mounting competition. Beyond NVIDIA, Qualcomm and MediaTek are investing heavily in on-device AI, while Huawei’s Ascend lineup delivers full-stack domestic independent development. We look forward to MoreMoto’s upcoming new hardware releases.
Overall, MoreMoto Intelligence delivered far more than a single chip at WAIC 2026; it unveiled a comprehensive prototype ecosystem for on-device AI. Has MoreMoto secured its place at the table? Based on this WAIC showcase, Leitech believes the company is firmly in the game.
The WAIC 2026, themed “Intelligent Partners, Co-Creating the Future,” is underway!
The AI narrative is shifting—from stacking model parameters toward practical Agent-driven productivity. Heterogeneous computing and photonic computing continue pushing computational ceilings upward. Embodied intelligence accelerates real-world applications, bringing physical AI to life as robots enter homes and factories.
The LeiTech WAIC reporting team has arrived in Shanghai to capture the annual pinnacle of AI industrialization—stay tuned!


雷科技







