HomeElectronics NewsA Free Installer Turns Gaming PC Into Local AI Servers 

A Free Installer Turns Gaming PC Into Local AI Servers 

A free open-source tool spreads a 125-billion-parameter AI model across GPU memory, system RAM and SSD storage, bringing private AI inference to consumer gaming PCs.

A voxel-art pagoda garden with cherry blossoms, generated by Qwen3.8-Flash-Next running locally through Strata on an RTX 5070
A voxel pagoda garden written by the AI model itself, running locally through Strata on a consumer RTX 5070. (Image: Niko1221 / Strata GitHub repository)

A gaming PC with a 12 GB graphics card can now run a 125-billion-parameter AI model locally, without an internet connection, cloud subscription or dedicated server. The open-source project Strata achieves this by distributing the model across the graphics card, system memory and SSD, allowing hardware that would normally be considered too small for such a model to handle local inference.

Developed by GitHub user Niko1221, Strata provides a browser-based chat interface as well as local API endpoints that can be connected to compatible AI applications. The project was first released in September 2026 and is designed to make large-model inference practical on consumer hardware.

The system is built around a mixture-of-experts model. Instead of activating the entire 125-billion-parameter model for every token, only a small group of its 24,576 expert components is used at a time. Strata distributes these components according to how frequently they are needed. Frequently accessed data is kept in GPU memory, while other parts remain in system RAM or are retrieved from the SSD.

A roughly 29 GB lookup table is stored on the SSD and accessed as required during generation. This arrangement allows the system to use a graphics card with considerably less VRAM than would normally be required to hold the complete model.

Strata also uses speculative decoding to improve generation speed. A smaller predictor proposes upcoming tokens, which the main model can then verify together rather than processing every token independently. According to the project, this can provide a 1.6× to 1.8× speed improvement without changing the final output. The underlying inference engine uses the open-source llama.cpp and ggml projects.

The hardware requirement is still substantial. Strata supports NVIDIA GeForce RTX 30, 40 and 50 series GPUs and selected AMD Radeon RX 7800, 7900, 9060 and 9070 series cards with at least 12 GB of VRAM. The project recommends 64 GB of system RAM, while its documentation gives different memory requirements depending on the selected model and compression level. An SSD with roughly 80 GB of free space is also recommended.

Installation is intended to be straightforward. Windows users can download the repository, open the project folder and run its setup script, while Linux users can run the corresponding shell script. The installer detects the available hardware and asks users to select a model, compression level, context size and image-input option before downloading the required files.

Depending on the selected compression, the download is approximately 66 GB to 76 GB. Once installed, users can access the local chat interface through a web browser. Strata also provides OpenAI-compatible and Anthropic-compatible API endpoints, allowing supported coding assistants and other AI applications to connect to the local model rather than a cloud service.

The project’s published benchmarks show what this approach can deliver. On an RTX 5070 with 12 GB of VRAM, the Q2_0 configuration reaches about 95 tokens per second, while the IQ2_XS and IQ3_XXS configurations reach about 78 and 66 tokens per second, respectively. With a 128K-token context, the reported generation speeds fall to 65, 52 and 45 tokens per second.

The trade-off is hardware and storage. The first launch can load tens of gigabytes of data into system memory and may temporarily make the computer unresponsive. Long conversations also increase processing requirements because the system has to handle a growing context. Strata is therefore better suited to a single user than a shared AI server.

For Indian users, the project offers an interesting way to repurpose an existing gaming PC into a private AI workstation. Prompts and responses can remain on the local machine, which can be useful when working with sensitive documents or code. There is also no recurring AI subscription fee.

The catch is the upfront hardware requirement and a download approaching 80 GB. For users who already own a gaming PC with sufficient VRAM and RAM, however, Strata demonstrates how consumer hardware can move large-scale AI inference closer to the desktop without depending on a remote cloud service.

For more information, click here. 

Loading form…
Ananthu Ashok
Ananthu Ashok
Ananthu Ashok is a tech journalist and has a deep interest in embedded systems, open source, IoT, robotics and emerging tech.

SHARE YOUR THOUGHTS & COMMENTS

EFY Prime

Unique DIY Projects

Electronics News

Truly Innovative Electronics

Latest DIY Videos

Electronics Components

Electronics Jobs

Calculators For Electronics