A 312,000-parameter language model running on an ESP32-S3 converts plain-English instructions into GPIO actions without Wi-Fi, cloud services or an external API.

Type “blink pin 4 every 5 seconds” into a serial terminal and an ESP32-S3 can make the connected LED blink. The board can also understand commands such as naming a pin “desk lamp”, switching several pins together, or refusing an invalid GPIO or unsafe timing value.
The open-source esp32-gpio-llm project does this by running a 312,000-parameter transformer directly on the microcontroller. The model is designed specifically for GPIO control rather than general conversation, keeping the entire process offline.
The model contains 312,000 parameters and occupies about 1.2 MB of flash in FP32 format. The ESP32-S3 maps the model and reads it from flash instead of requiring the complete model to sit in internal RAM.
During inference, the device uses about 300 KB of PSRAM for its key-value cache. The repository reports command latency between 150 milliseconds and 1.5 seconds, measured on the ESP32-S3.
The model was trained from scratch on synthetic GPIO-control commands. Its output is converted into a small command structure rather than free-form JSON, limiting what can reach the hardware layer.
The project does not rely on the language model alone to decide whether a GPIO operation is safe. A separate hardware layer validates the requested pin and parameters.
The current implementation permits 25 GPIOs: 1–18, 21, 38–42 and 48. It also restricts blink intervals to 50–10,000 milliseconds. For example, a request to blink a pin every 60 seconds is rejected instead of being executed.
This separation became important during development. The project found that making illegal values impossible for the model to represent could cause a different problem: an invalid request such as “pin 100” could be transformed into a valid pin number. The current approach allows the model to reproduce the requested value while leaving the hardware layer responsible for rejecting it.
The firmware can associate a GPIO with a user-defined name. A user can type “call pin 4 the desk lamp” and subsequently use “turn on the desk lamp.”
Names are stored by the device and persist across a power cycle. Unknown names are rejected rather than being mapped to the nearest known device.
For a basic demonstration, an LED and a resistor can be connected between a supported GPIO and ground. The repository recommends an ESP32-S3 with PSRAM; its quick-start firmware can be flashed directly from the project’s releases.
The latest evaluation used 612 valid commands and 417 commands that should be refused, tested across three training seeds.
The model achieved 84.4 per cent ± 1.5 per cent exact-match accuracy. It moved the wrong pin on 4.5 per cent ± 0.7 per cent of valid commands and accepted input that should have been rejected at 5.1 per cent ± 1.3 per cent when measured by whether a pin actually moved. The broader false-accept metric was 13.3 per cent ± 1.6 per cent, while pin-copy accuracy reached 91.3 per cent ± 0.9 per cent.
These figures remain below the project’s own targets of more than 95 per cent exact match, less than 2 per cent false accepts, more than 99 per cent pin copying and zero substitutions.
Interestingly, the repository also compares the model with its conventional rule-based parser. The parser correctly handled 41.7 per cent of the held-out commands but did not perform incorrect GPIO actions, because it simply rejected commands outside its predefined grammar. The language model provides substantially broader wording coverage, but introduces the possibility of incorrect actions.
The project is not intended to replace a chatbot or provide general-purpose AI. It demonstrates how a narrowly trained transformer can translate natural-language instructions into physical hardware actions while remaining completely offline.
Its MIT-licensed source includes the training pipeline, firmware and supporting data, making the complete path from model training to GPIO control available for experimentation.
For makers and engineering students, the more significant lesson is architectural: a microcontroller does not need a cloud connection to interpret natural-language commands, but the model should not be treated as the final safety boundary when it controls physical hardware.
For more information, click here.



