Hook
LLM inference in C/C++. Contribute to ggml-org/llama.cpp development by creating an account on GitHub.
More breakout videos from this creator.
Run AI model from USB, fully offline. First, grab a USB drive with at least 8GB of storage and plug it into your laptop. Right-click the drive in File Explorer, select Format, choose exFAT, enable Quick Format, and click Start. Once it's formatted, go to Google and search "llamafile GitHub." Open the official result, go to Releases, select version 0.10.5, and download it. Rename the file so it ends with .exe, then copy it onto your USB drive. Next, search "hugging face," open the website, and go to Models. Under Libraries, select GGUF, and search for "qwen." Open the model, go to Files and versions, and download the .gguf file that's around 2.5 GB. Copy that file to your USB drive as well. Now open Notepad and create a command using the executable file name, server, model, and the .gguf file name. Save it as start.bat directly on your USB drive. Plug the USB drive into another laptop, open the drive, and double-click start.bat. A terminal will open and display the local IP address. Open that address in your browser, and your personal AI chat interface will be ready. You can now run the AI locally without sending your prompts to a cloud service. Follow for more AI and tech tips.