百科.dev
登录
> 返回资讯列表
news_article.exe
📰
#GPT#Claude

如何以 VS 代码、 打开代码和 LM 工作室本地运行自由 AI 编码助理

How to Run a Free AI Coding Assistant Locally with VS Code, opencode, and LM Studio

2026年9月6日1 次浏览来源:Dev.to 阅读原文

每个AI编码助理都想要你的信用卡 还有你的密码 还有第三种选择,不费钱:用你自己的电脑运行整个程序。 没有订阅。 没有互联网。 你的密码从不离开机器 我用一个很普通的电脑设计了这个,它有用。 就是这样 你们正在建造的三部作品一起工作: LM Studio ——运行你的PC上的AI模型. 这是引擎。 opencode — 读取和编辑您文件的编码助手。 VS代码——你写代码的地方。 LM工作室负责思考 打开代码可以工作。 维基百科 密码是你坐的地方 我的设置:Windows,RTX 3070 Ti 有8GB VRAM,16GB RAM. 这是一个中程游戏PC,不是工作站。 如果你的相似,你没事。 第一步:从lmstudio.ai安装LM Studio下载并安装...

Every AI coding assistant wants your credit card. And your code. There's a third option, and it costs nothing: run the whole thing on your own computer. No subscription. No internet. Your code never leaves your machine. I set this up on a fairly ordinary PC, and it works. Here's exactly how. What you're building Three pieces working together: LM Studio — runs the AI model on your PC. This is the engine. opencode — the coding assistant that reads and edits your files. VS Code — where you actually write code. LM Studio does the thinking. opencode does the work. VS Code is where you sit. My setup: Windows, RTX 3070 Ti with 8GB VRAM, 16GB RAM. That's a mid-range gaming PC, not a workstation. If yours is similar, you're fine. Step 1: Install LM Studio Download it from lmstudio.ai and install it like any normal app. LM Studio lets you download and run open-source AI models directly on your computer. It's the easiest way into local AI — no command line required. Step 2: Download the model Open the search inside LM Studio and look for qwen3 8b. You'll see a lot of versions. Check these three things before you download: Format: GGUF Quantization: Q4_K_M Capabilities: "tool use" must be listed Why those matter, in plain English: Quantization is compression. Q4_K_M shrinks the model so it fits on a smaller graphics card. You lose a little quality, but you gain the ability to actually run it. Tool use is non-negotiable. A coding assistant needs to open your files and edit them. A model without tool use can only chat about your code — it can't touch it. Skip this check and nothing will work later. Got different hardware? Pick a model that fits it. Bigger models are smarter but hungrier. An 8B model is a comfortable fit for 8GB of VRAM. Once the download finishes, your model shows up under My Models. Step 3: Load the model Go to the Developer tab and select your model. Turn on "Manually choose model load parameters", then click the small arrow next to the model name to open the settings. Now find context size and set it to 16000. Context size is how much text the model can hold in its head at once — your question plus its answer plus whatever code it's looking at. Bigger context means it understands more of your project. It also eats more VRAM. 16000 is a good number for 8GB. If you have less, go lower. If the model refuses to load, lower it again and try once more. Step 4: Turn on the server Flip the server status to Running. Your model is now live at . That address is your own machine talking to itself. Nothing is going out to the internet — which is the entire point. Step 5: Install VS Code and opencode Install VS Code if you don't have it. Then install opencode. On Windows, the simplest route is npm: Open a terminal inside VS Code (Terminal → New Terminal) and type: opencode starts up right there in the terminal panel. Step 6: Point opencode at your local model Here's the part that trips people up. opencode has no idea your model exists yet. You have to tell it, using a config file. Here's the config I use: Three lines matter here: points at LM Studio on your own machine. Same address from Step 4. can be anything. LM Studio doesn't check it, but opencode expects the field to exist. is what lets the assistant edit your files. Remember the tool use check back in Step 2? This is where it pays off. Save the file to: Swap for your actual Windows username. If that folder doesn't exist, create it. You can also grab it from this gist. Step 7: Pick your model and test it Restart VS Code, then start opencode again. Type and select qwen/qwen3-8b from the list. Now ask it to do something real. Give it a file to fix. Want proof it's actually running locally? Switch over to LM Studio and check the logs. You'll see tokens streaming as the assistant types. That's your own GPU doing the work. What to expect Let's be straight about this: a local 8B model is not going to match Claude or GPT on a hard architectural problem. It's smaller, and smaller means less capable. But for the everyday stuff — writing boilerplate, explaining unfamiliar code, catching bugs, renaming things across files — it holds its own. And it's fast, because there's no network round trip. The best part is what it costs: nothing, forever. No token limits. No monthly bill. Works on a plane. If you have more VRAM than I do, try a larger model. Same steps, better results.

> 分享: