HomeElectronics NewsGoogle Toolkit Brings Local AI To Raspberry Pi 5

Google Toolkit Brings Local AI To Raspberry Pi 5

Google’s LiteRT toolkit lets a Raspberry Pi 5 run compact Gemma language and vision models locally, without sending inference to a cloud service.

Gemma translator
Gemma translator

A tutorial published on Raspberry Pi’s website demonstrates running Google’s Gemma family of AI models on-device using LiteRT on a Raspberry Pi 5. Written by Naush Patuck, a Raspberry Pi Software Engineering Manager, the walkthrough pairs LiteRT with the Reachy Mini robot to process vision and language tasks locally rather than sending them to a cloud service. 

- Advertisement -

LiteRT is Google’s on-device runtime for deploying trained AI models on local hardware. The Raspberry Pi tutorial demonstrates running several Gemma models locally on a Raspberry Pi 5, including the compact Gemma 3 270M and EmbeddingGemma 300M models, as well as the larger Gemma 4 E2B and E4B variants. The workflow also uses YOLO26n for object detection and MediaPipe Selfie Segmenter for image-processing tasks. 

On the 8 GB Raspberry Pi 5, Gemma 3 270M achieved 433.17 prefill tokens per second and 22.58 decode tokens per second using LiteRT-LM. The larger Gemma 4 E2B model reached 99 prefill tokens per second and 9 decode tokens per second, with an end-to-end generation speed of about 27.3 characters per second. For computer vision, YOLO26n recorded a latency of 101.26 ms on the CPU and 375.73 ms using LiteRT’s WebGPU Vulkan backend on the GPU. 

Running Gemma models locally on a Raspberry Pi 5 avoids sending inference requests to a cloud service and allows AI processing to take place without an internet connection. The trade-off is lower inference performance compared with more powerful computing platforms. The 8 GB Raspberry Pi 5 currently carries a $95 official list price, while Indian retail prices vary by seller. 

- Advertisement -

The LiteRT command-line tool download is about 25 MB, while the Gemma models vary considerably in size, from compact models such as Gemma 3 270M and EmbeddingGemma 300M to multi-gigabyte variants. Storage and memory requirements therefore need to be considered when running different models on a Raspberry Pi 5. The larger models also deliver lower generation rates, making model selection important for interactive applications.

Local inference on a Raspberry Pi 5 gives developers a way to experiment with generative AI on-device without relying on a cloud connection. The combination of LiteRT and compact Gemma models provides a practical platform for testing AI applications where local processing is preferred.

Loading form…
Ananthu Ashok
Ananthu Ashok
Ananthu Ashok is a tech journalist and has a deep interest in embedded systems, open source, IoT, robotics and emerging tech.

SHARE YOUR THOUGHTS & COMMENTS

EFY Prime

Unique DIY Projects

Electronics News

Truly Innovative Electronics

Latest DIY Videos

Electronics Components

Electronics Jobs

Calculators For Electronics