A technical proposal for local AI

Load the model
for the work.

SkillWeave is a benchmark-guided layer that prepares a shared language model for a workload before inference begins—allocating precision and memory across disk, RAM, and VRAM.

See an example

The thesis

One shared model.
A deliberate working set.

General-purpose models carry useful capacity that a local agent may rarely need. Conventional loaders nevertheless materialize a broad representation with little regard for the task in front of them.

SkillWeave makes that loading step workload-aware. It chooses which weight groups are available at which precision and in which memory tier—without claiming that a dense transformer can simply discard every non-salient parameter.

01Requested
workload
→
02Materialization
plan
→
03Disk · RAM ·
VRAM
→
04Original model
runtime

The system

Specialization begins with representation and residency—not a new model per skill.

01

Preserve the execution contract

SkillWeave loads an existing model through a different representation and residency policy. Its forward architecture remains unchanged unless a separately validated compiled variant is requested.

02

Allocate a budget, not a mythology

Dense models do not contain a small removable set of domain weights. The system assigns finite precision and memory budgets to groups that matter for a measured workload.

03

Measure the whole hierarchy

Quality is evaluated alongside cold start, disk reads, RAM and VRAM pressure, cache behavior, page-fault stalls, and steady-state throughput.

Materialization plan

Every group is somewhere.

For dense models, all required groups remain available through a low-bit base. Selected groups may receive higher-fidelity overlays or fast-memory residency. For MoE models, likely experts can be promoted into cache while others remain compressed or paged.

01
DiskCompressed fallback

Low-bit random-access pages

02
RAMMapped & cached

Hot pages and async prefetch

03
VRAMSelected working set

High-fidelity overlays where they count

A concrete example

Repair a broken
frontend build.

A repository-level task gives the small profiler enough signal to select a starting plan—before the large model enters VRAM.

01Request + repository

“Fix the React build and preserve accessibility.”

The profiler sees TypeScript, React, package metadata, and a failed build—not a vague category named “JavaScript.”

→
02Load plan

Keep the base compact. Promote the relevant working set.

Low-bit base pages remain available; selected representation overlays and likely MoE experts receive faster residency.

→
03Unchanged execution

Generate, inspect, edit, then verify.

The original model runtime works with retrieval and tools. Tests, type checks, and browser checks decide whether the answer is good.

Important

This is not a claim that only “React weights” are loaded. Dense computation still requires every group to be available somewhere; the plan controls representation and residency.

What must be proven

No promises
without measurement.

Hidden-task successCold-start timeDisk bytes readPeak VRAM & RAMCache-hit ratePage-fault stallsTool correctnessGeneral-task regression

MVP roadmap

Start with a package, planner, materializer, and benchmark harness—not a speculative giant model.

  1. 01

    Materialize

    Package a low-bit random-access base and compare workload-guided loading against conventional loading.

  2. 02

    Allocate

    Validate trace-guided precision and residency maps against uniform quantization at the same device budget.

  3. 03

    Specialize

    Add adapters, pruning, or distillation only where measured gains justify the extra artifact.

Read the proposal

Specialization should be
observable.

The full paper defines the architecture, evaluation gates, memory realities, and research risks behind SkillWeave.