Preserve the execution contract
SkillWeave loads an existing model through a different representation and residency policy. Its forward architecture remains unchanged unless a separately validated compiled variant is requested.
A technical proposal for local AI
SkillWeave is a benchmark-guided layer that prepares a shared language model for a workload before inference begins—allocating precision and memory across disk, RAM, and VRAM.
The thesis
General-purpose models carry useful capacity that a local agent may rarely need. Conventional loaders nevertheless materialize a broad representation with little regard for the task in front of them.
SkillWeave makes that loading step workload-aware. It chooses which weight groups are available at which precision and in which memory tier—without claiming that a dense transformer can simply discard every non-salient parameter.
The system
Specialization begins with representation and residency—not a new model per skill.
SkillWeave loads an existing model through a different representation and residency policy. Its forward architecture remains unchanged unless a separately validated compiled variant is requested.
Dense models do not contain a small removable set of domain weights. The system assigns finite precision and memory budgets to groups that matter for a measured workload.
Quality is evaluated alongside cold start, disk reads, RAM and VRAM pressure, cache behavior, page-fault stalls, and steady-state throughput.
Materialization plan
For dense models, all required groups remain available through a low-bit base. Selected groups may receive higher-fidelity overlays or fast-memory residency. For MoE models, likely experts can be promoted into cache while others remain compressed or paged.
Low-bit random-access pages
Hot pages and async prefetch
High-fidelity overlays where they count
A concrete example
A repository-level task gives the small profiler enough signal to select a starting plan—before the large model enters VRAM.
The profiler sees TypeScript, React, package metadata, and a failed build—not a vague category named “JavaScript.”
Low-bit base pages remain available; selected representation overlays and likely MoE experts receive faster residency.
The original model runtime works with retrieval and tools. Tests, type checks, and browser checks decide whether the answer is good.
This is not a claim that only “React weights” are loaded. Dense computation still requires every group to be available somewhere; the plan controls representation and residency.
What must be proven
MVP roadmap
Start with a package, planner, materializer, and benchmark harness—not a speculative giant model.
Package a low-bit random-access base and compare workload-guided loading against conventional loading.
Validate trace-guided precision and residency maps against uniform quantization at the same device budget.
Add adapters, pruning, or distillation only where measured gains justify the extra artifact.
Read the proposal
The full paper defines the architecture, evaluation gates, memory realities, and research risks behind SkillWeave.