<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>machine-learning on Machine Learning Stuffs</title><link>https://moisesvw.github.io/categories/machine-learning/</link><description>Recent content in machine-learning on Machine Learning Stuffs</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Sat, 05 Sep 2026 16:30:00 -0500</lastBuildDate><atom:link href="https://moisesvw.github.io/categories/machine-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>Running Qwen3.8-27B on a 6 GB Laptop GPU (Slowly, but for Real)</title><link>https://moisesvw.github.io/2026/september/qwen38-27b-on-6gb-laptop/</link><pubDate>Sat, 05 Sep 2026 16:30:00 -0500</pubDate><guid>https://moisesvw.github.io/2026/september/qwen38-27b-on-6gb-laptop/</guid><description>This month my feeds filled up with posts about Qwen3.8-27B running at 50 to 75 tokens per second on an RTX 3090 with llama.cpp. Great numbers. My spare Linux laptop has an RTX 3060 Laptop GPU with 6 GB of VRAM, which is not even half of what the Q4 weights need. I wanted to know if the model could run on it at all.
It can. As I write this, a coding agent on my LAN is talking to Qwen3.</description></item></channel></rss>