How to Use GLM-5.3 Flash for Free in VS Code (Kilo Code)

October 3, 2026 | Umair Alam | 8 min read

The era of paying $20 to $40 every month for proprietary AI coding subscriptions like Cursor Pro, GitHub Copilot Enterprise, or Windsurf is officially ending. Frontier AI labs and model aggregators are making world-class reasoning and coding models available for free or at near-zero rates. Among the most impressive breakthroughs in agentic software engineering is GLM-5.3 Flash—a state-of-the-art Mixture-of-Experts (MoE) model built with a massive 1-million-token context window, lightning-fast generation speeds, and native multi-tool agent capabilities.

TL;DR: You can run GLM-5.3 Flash completely free inside Visual Studio Code by combining B.AI (b.ai) with the open-source Kilo Code extension. B.AI provides generous free registration credits and an OpenAI-compatible REST API (https://api.b.ai/v1). By plugging your free B.AI API key into Kilo Code, you get an autonomous, multi-file coding agent with full terminal execution and project-wide reasoning—without spending a dime on monthly subscriptions.


What You’ll Learn

  • Why GLM-5.3 Flash is a formidable, free alternative to Claude 3.5 Sonnet and GPT-4o
  • Architectural specifications: 320B MoE parameters, 1M context window, and tool calling
  • How to register on B.AI (b.ai) and claim free developer testing credits
  • How to install and configure the Kilo Code agentic extension in Visual Studio Code
  • How to connect Kilo Code to B.AI using standard OpenAI-compatible endpoints (https://api.b.ai/v1)
  • Hands-on testing: autonomous project scaffolding, multi-file refactoring, and terminal test verification
  • Optimization tips: managing the 1M context window, temperature tuning, and rate limits

GLM-5.3 Flash Architecture & Technical Specs

Before setting up your IDE, here is how GLM-5.3 Flash compares to leading commercial alternatives:

Specification GLM-5.3 Flash (via B.AI) Claude 3.5 Sonnet DeepSeek V3
Underlying Model Family GLM Series (Zhipu / Z.ai) Anthropic DeepSeek AI
API Provider B.AI (b.ai) Anthropic Console DeepSeek Platform
Model Architecture Mixture-of-Experts (320B Total / 18B Active) Dense / Speculative MoE Mixture-of-Experts (MoE)
Context Window 1,000,000 Tokens (1M) 200,000 Tokens (200K) 64,000 Tokens (64K)
Native API Protocol OpenAI-Compatible (https://api.b.ai/v1) Anthropic Messages API OpenAI-Compatible (/v1)
VS Code Agent Harness Kilo Code Claude Code CLI Third-party IDE extensions
Free Access Tier Yes (Free Registration Credits on B.AI) Limited Free Web / $20/mo or Pay-Per-Token Limited Free Web / Low Cost
Primary Sweet Spot High-Speed Agentic Tool Use & Large Codebases Nuanced Reasoning Cost-Effective Cloud Reasoning

Why Pair GLM-5.3 Flash with Kilo Code in VS Code?

Most developers prefer not to migrate their entire workflow to locked-down, proprietary editors like Cursor. Pairing GLM-5.3 Flash with Kilo Code inside standard Visual Studio Code offers several massive advantages:

  1. Keep Your Existing Setup: You retain all your custom keybindings, themes, formatters, and Git configurations inside standard VS Code.
  2. Massive 1M Context Window: While typical coding assistants choke when scanning large monorepos, GLM-5.3 Flash can ingest your entire codebase, API documentation, and dependency graph simultaneously without dropping instructions.
  3. Blistering Token Velocity: “Flash” models are optimized for speed (often 60–100+ tokens per second). When an autonomous agent runs multi-step loops (reading files, planning diffs, executing tests, debugging syntax), speed is the difference between an effortless 30-second fix and a frustrating wait.
  4. Autonomous Agent Execution: Unlike simple autocomplete plugins, Kilo Code can create files, modify existing code across multiple folders, search directories, and execute shell commands directly in your VS Code terminal.

Step 1: Claim Free GLM-5.3 Flash API Credits on B.AI

To use GLM-5.3 Flash in VS Code without paying upfront, generate a free API key on the B.AI platform:

  1. Visit the official B.AI platform at b.ai.
  2. Click Sign Up / Register to create a free account using your email or Google/GitHub login.
  3. Upon registration, B.AI provides free bonus credits to test frontier models on the platform.
  4. Navigate to the API Keys section in your B.AI dashboard.
  5. Click Create API Key, name it VSCode-KiloCode-GLM, and copy your generated key to your clipboard.
  6. Verify the B.AI connection parameters:
    • Base URL: https://api.b.ai/v1
    • Model Name: glm-5.3-flash (or glm-5.3-flashx for the high-speed variant)

Security Tip: Store your API key securely. Never commit API keys to public repositories or hardcode them into shared codebases.


Step 2: Install Kilo Code in Visual Studio Code

Kilo Code is an open-source, agentic AI coding assistant built specifically for VS Code:

  1. Open Visual Studio Code.
  2. Open the Extensions sidebar by pressing Ctrl + Shift + X (Windows/Linux) or Cmd + Shift + X (macOS).
  3. Search for Kilo Code.
  4. Click Install.
  5. Once installed, a new Kilo Code icon will appear on your left Activity Bar.

Step 3: Configure B.AI & GLM-5.3 Flash in Kilo Code

Now connect Kilo Code to B.AI’s OpenAI-compatible endpoint:

  1. Click the Kilo Code icon in the left Activity Bar to open the extension panel.
  2. Click the Settings (Gear) icon in the top-right corner of the Kilo Code panel.
  3. Under Provider, select Custom Provider and click on the Connect button.
  4. Enter the B.AI endpoint details:
    • Provider ID:
      bai
    • Display Name:
      B-AI / any name of your choice
    • Provide API:
      OpenAI Compatible
    • Base URL:
      https://api.b.ai/v1
    • API Key: Paste your B.AI API key from Step 1.
    • Model ID:
      glm-5.3-flash
  5. Click Submit to save the settings.

Step 4: Understanding Kilo Code Modes

Kilo Code provides distinct modes to tackle different developer workflows:

  • Code Mode: The default autonomous execution mode. Kilo Code actively inspects your file structure, creates new files, modifies existing code with precise diffs, and runs terminal commands to test its work.
  • Ask Mode: A read-only conversational mode for explaining code snippets, understanding legacy logic, or asking architectural questions without modifying files.
  • Debug Mode: Specifically designed for problem-solving. You can paste error logs, stack traces, or compiler messages, and the agent will trace the underlying code logic to propose direct bug fixes.
  • Plan Mode: A technical planner used to map out system designs, feature scope, and software architecture before coding.

Step 5: Testing Autonomous Project Execution

To test your new setup, open an empty folder or existing project in VS Code and give Kilo Code an autonomous task in Coder Mode:

Inspect this project. Build a production-ready Node.js Express microservice in `server.js` with:
1. A health check route at `/api/health`
2. In-memory rate limiting middleware
3. Structured JSON request logging
4. A complete test suite using Jest in `test/server.test.js`
Install necessary dependencies, run the test suite in the terminal, and make sure all tests pass.

Watch how the agent executes the task:

  1. Reads Workspace: Analyzes existing files and determines whether package.json exists.
  2. Scaffolds Files: Creates server.js and test/server.test.js with clean, modular code.
  3. Executes Terminal Commands: Runs npm install express jest supertest, launches the test runner, and reads the output directly from your terminal.
  4. Self-Heals: If a test fails, Kilo Code reads the stack trace, modifies the code, and reruns the test until it passes.
  5. All of this executes rapidly thanks to GLM-5.3 Flash’s high token generation speed.

Best Practices & Optimization Tips

  • Leverage the 1M Context Window: Don’t hesitate to reference full documentation files or ask Kilo Code to “Scan all files in `src/` to find unused utility functions.” GLM-5.3 Flash’s 1M context easily handles large multi-thousand-line repositories.
  • Review Shell Commands: Kilo Code prompts for permission before executing sensitive terminal commands. Always review rm, database migration, or deployment commands before approving execution.

Key Takeaways & What to Explore Next

Pairing GLM-5.3 Flash via B.AI with Kilo Code gives you an enterprise-grade AI coding agent completely free:

  • Zero Monthly Subscriptions: Avoid the $20/month Cursor Pro or Copilot tax.
  • Full Autonomy: Multi-file edits, project-wide understanding, and terminal execution inside your native VS Code environment.
  • Massive 1M Context: Analyze entire repositories without losing track of instructions.

If you are expanding your free developer toolkit, check out my companion tutorials:


Frequently Asked Questions (FAQ)

How do I get free access to GLM-5.3 Flash?

You can access GLM-5.3 Flash through B.AI (b.ai), which offers free registration bonus credits for developer testing. This allows you to generate an API key and connect it to VS Code without upfront subscription charges.

Why use Kilo Code instead of standard autocomplete extensions?

Traditional extensions (like basic Copilot autocomplete) only suggest the next few lines of code. Kilo Code is an autonomous agent—it reads your whole directory, writes code across multiple files, executes terminal commands, and debugs errors automatically.

Can I use other models with Kilo Code?

Yes. Kilo Code supports OpenAI-compatible endpoints, Anthropic, and local endpoints (like Ollama or LM Studio), allowing you to switch between cloud models and local offline agents effortlessly.

Umair Alam

Written by

Umair Alam

Business Automation Specialist & Technical Developer with 17+ years of experience in enterprise IT, database development, and modern web technologies. I specialize in bridging the gap between complex organizational workflows and automated digital systems.