Open Source Code Completion Models for VS Code
Relying on proprietary cloud extensions for code completion creates an uncomfortable architectural trade-off for engineering organizations. While tools like GitHub Copilot accelerate development velocity, they route proprietary source code, internal APIs, and cryptographic secrets through third-party cloud infrastructure. Replacing closed services with self-hosted, open-weight models allows engineering teams to maintain complete control over their intellectual property while eliminating data exfiltration risks.
Architectural Separation: Client Extensions vs Inference Backends
Implementing a self-hosted completion setup requires decoupling the client-side IDE interface from the backend inference engine. The local IDE extension captures the active cursor context, extracts the surrounding prefix and suffix code buffers, and transmits a payload to a remote inference server. The server executes the model forward pass and streams generated tokens back to the editor buffer in real time.
For European organizations, this decoupling establishes a crucial operational boundary. Under the US CLOUD Act, US-based providers can be compelled to produce data in their possession, custody, or control regardless of where that data is physically stored, so choosing a European cloud region does not by itself remove the exposure. By self-hosting open-weight models on sovereign infrastructure, engineering teams ensure that proprietary source code never leaves European jurisdiction.
- Client-side IDE plugins handle cursor telemetry, local file heuristics, and Fill-in-the-Middle context formatting.
- Remote high-performance inference engines host open-weight model weights on dedicated or serverless GPU compute.
- Low-latency HTTP transport layers stream tokens back to the editor without crossing transatlantic network boundaries.
Understanding the Codex CLI Open Source Status
Navigating the OpenAI developer ecosystem reveals a sharp divide between command-line tooling and IDE integrations. While OpenAI released the Codex CLI under the permissive Apache-2.0 license, the core extensions embedded within developer IDEs have remained proprietary, closed-source binaries.
The Opacity of Closed IDE Extensions
A closed IDE plugin acts as an opaque black box inside your developer environment. It reads workspace files, intercepts editor keystrokes, and packages contextual code snippets according to closed heuristics. For security-conscious teams, this lack of transparency prevents internal security audits from verifying what telemetry, metadata, or source files are transmitted over the wire.
This architectural opacity has led to persistent community demand for open-source client implementations. Developers and security leads repeatedly ask for auditable extension code so they can verify what is transmitted over the wire, and for the ability to point their editors at self-hosted backends instead of vendor-managed endpoints.
Can You Use Codex with Open Source Models?
Official Codex workflows and extensions are tightly coupled to vendor-managed endpoints, preventing developers from re-pointing them to custom local or private cloud servers. However, engineering teams can bypass this vendor lock-in entirely by using open-source inference engines that expose standard API interfaces.
OpenAI API Compatibility via Modern Inference Engines
Modern open-source serving frameworks act as drop-in replacements for proprietary APIs. Serving engines like vLLM provide a built-in OpenAI-compatible server that implements standard completions and chat endpoints. This standardization allows any client tool designed for OpenAI protocols to communicate with open-weight models without code modifications.
Deploying an OpenAI-compatible inference engine allows teams to serve state-of-the-art open-weight models such as Qwen or DeepSeek behind a standard HTTP interface. The inference engine manages tokenization, continuous batching, and KV cache memory allocation while remaining completely transparent to the client IDE.
- Deploy an inference engine like vLLM or TensorRT-LLM that implements the OpenAI /v1/completions specification.
- Load an open-weight code model supporting Fill-in-the-Middle token formatting into GPU memory.
- Configure your IDE client with the custom base URL and API bearer token pointing to the self-hosted endpoint.
VS Code and JetBrains Extension Options
Transitioning away from closed AI assistants requires selecting IDE extensions engineered for configurable backends. Both VS Code and JetBrains environments support flexible plugins that route inline completions and chat requests to arbitrary endpoints.
Continue.dev: Open-Source Multi-IDE Integration
Continue.dev has emerged as the standard open-source AI coding assistant across VS Code and JetBrains IDEs. It allows developers to specify custom model providers, modify context aggregation heuristics, and configure distinct endpoints for inline autocompletion and conversational chat.
JetBrains AI Assistant Custom Provider Support
JetBrains native AI Assistant also supports custom OpenAI-compatible endpoints. Under Providers & API keys you can pick an OpenAI-compatible third-party provider for chat and core features, and in the AI Completion section select the "OpenAI Compatible" provider for inline code completion and next edit suggestions, supplying an API key, a base URL, and the model name. The same settings page exposes a prompt schema selector, which JetBrains sets automatically for recognised models and which you override manually when the served model is not recognised.
| Extension | Supported IDEs | License | Custom Base URL Support | FIM Autocomplete Support |
|---|---|---|---|---|
| Continue.dev | VS Code, JetBrains, Neovim | Apache-2.0 | Native configuration via config.yaml / config.json | Yes (via FIM-capable open models) |
| JetBrains AI Assistant | JetBrains IDEs (IntelliJ, PyCharm, etc.) | Proprietary | Configurable under Providers & API Keys | Yes (requires FIM prompt schema) |
| Refact | VS Code, JetBrains | Open-Source (BSD-3-Clause) | Self-hosted server integration | Yes (specialized FIM architecture) |
Setting Up an OpenAI-Compatible Endpoint Server
Deploying a production-grade completion backend begins with containerizing an optimized inference engine. Running your serving stack via Docker with the NVIDIA Container Toolkit exposes host GPUs to the container and guarantees consistent CUDA kernel execution across on-demand GPU instances, packaging model weights, the inference server, and CUDA dependencies into one reproducible unit.
Configuring Continue.dev for Custom Endpoints
Once your backend inference engine is running, connecting Continue.dev requires updating its local configuration file (config.yaml or config.json). Continue's self-hosting guide states that when the API you use is OpenAI-compatible, you can use the "openai" provider and change the baseUrl to point to your server; basic authentication works with any provider through the apiKey field, which translates to an Authorization Bearer header, while custom auth headers and client certificates are passed through requestOptions.
The configuration below demonstrates how to configure Continue.dev to route inline tab completions to a remote OpenAI-compatible server:
- tabAutocompleteModel: Defines the dedicated model configuration used exclusively for inline code autocompletion.
- title: A descriptive label for the model displayed within the IDE interface.
- provider: Configured as 'openai' to use the standard OpenAI completions API protocol.
- model: The exact model identifier loaded on the backend inference engine (for example, Qwen3-Coder-30B-A3B).
- apiBase: The HTTPS base URL pointing to the remote server (for example, https://api.lyceum.technology/v1).
- apiKey: The authentication bearer token passed in the Authorization header to secure the endpoint.
Separating the completion model configuration from interactive chat models ensures that lightweight, low-latency models handle keystroke-level completions while larger models handle multi-turn architectural discussions.
The Latency Budget for Inline Code Completion
Inline code completion operates under significantly stricter performance constraints than conversational chat workflows. While developers tolerate multi-second response times for chat responses, inline autocompletion must trigger and stream tokens almost instantaneously to prevent interrupting typing cadence.
Empirical Practitioner Latency Thresholds
Empirical research demonstrates that completion latency directly determines developer adoption. In an empirical study that interviewed 15 practitioners and surveyed 599 practitioners across 18 IT companies on code completion expectations, the authors report that latency is a key factor affecting the likelihood of adoption, and that a token-level completion tool returning candidates within 200 milliseconds satisfies 85% of practitioners.
Exceeding this 200 ms Time-To-First-Token (TTFT) threshold causes visual latency, missed keystrokes, and high suggestion rejection rates. Maintaining optimal inference latency requires minimizing network round-trip time (RTT), prefill computation over surrounding code context, and decode step duration.
- Network RTT: 15 ms to 35 ms for local intra-European routing, compared to 90 ms or more for transatlantic hops.
- Context extraction and tokenization: 10 ms to 20 ms in the local IDE background process.
- Prefill and Time-To-First-Token (TTFT): 80 ms to 120 ms on modern enterprise GPU hardware.
- Token decode generation: 15 ms to 25 ms per token for multi-token inline predictions.
Model Size Trade-Offs and Serverless Inference
Satisfying the 200 ms latency budget requires matching model parameter size to hardware throughput. Dense 70B and 405B models incur substantial prefill and decode latencies, making them inefficient choices for real-time keystroke autocompletion on standard instances.
Mixture-of-Experts and Dynamic Memory Management
Sparse Mixture-of-Experts (MoE) coding models provide the optimal balance between completion accuracy and execution speed. Models like Qwen3-Coder-30B-A3B activate only a fraction of their total parameters during each forward pass, drastically reducing memory bandwidth demands while preserving complex language and syntax comprehension.
At the serving layer, vLLM utilizes PagedAttention to partition the Key-Value (KV) cache into non-contiguous blocks, virtually eliminating memory fragmentation and reducing KV cache memory waste to near zero compared to traditional serving allocators. This dynamic allocation allows high concurrency without exhausting VRAM during spiky developer traffic.
For engineering teams that prefer to avoid managing GPU infrastructure directly, Lyceum Technology provides EU-hosted inference economics through Serverless Inference endpoints running in eu-north1. Teams can serve models such as Qwen3-Coder-30B-A3B with drop-in OpenAI SDK compatibility and per-token pricing, guaranteeing full data sovereignty under European jurisdiction without maintaining idle GPU clusters.
- Decouple your IDE by pairing open-source extensions like Continue.dev with OpenAI-compatible inference servers.
- Select sparse MoE coding models to stay strictly within the 200 ms Time-To-First-Token latency threshold.
- Utilize EU-native Serverless Inference to eliminate US CLOUD Act exposure while benefiting from predictable per-token compute costs.