Instructions to use rahul05ranjan/VIC-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use rahul05ranjan/VIC-2B with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("rahul05ranjan/VIC-2B") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use rahul05ranjan/VIC-2B with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "rahul05ranjan/VIC-2B"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "rahul05ranjan/VIC-2B" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use rahul05ranjan/VIC-2B with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "rahul05ranjan/VIC-2B"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "rahul05ranjan/VIC-2B" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rahul05ranjan/VIC-2B", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use rahul05ranjan/VIC-2B with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "rahul05ranjan/VIC-2B"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default rahul05ranjan/VIC-2B
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use rahul05ranjan/VIC-2B with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "rahul05ranjan/VIC-2B"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "rahul05ranjan/VIC-2B" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
VIC-2B
A 1.55 GB model that turns a question about a data file into a typed plan, which a deterministic engine then runs on your machine.
VIC (Vasuki Indicus Compute) is the planner model of vasuki. Given a request in plain words and a short profile of the tables it refers to, VIC-2B writes a plan of typed steps: filter, derive, aggregate, join, convert units and so on. The vasuki engine checks the plan against the columns and runs it. The model never computes a result, never writes code, and never sees the whole file.
| Developer | rahul05ranjan |
| Model type | Causal language model, fine-tuned to write vasuki plans as JSON |
| Base model | Qwen/Qwen3.5-2B, text part only (1.9B parameters) |
| This repository | MLX, 6-bit, 1.55 GB, for Apple silicon |
| Other formats | VIC-2B-GGUF: 8-bit GGUF, 2.0 GB, for Ollama and llama.cpp on any platform |
| Languages | English; Hinglish to a lesser degree |
| Licence | Apache 2.0 |
| Code | github.com/rahul05ranjan/vasuki-vic |
| Version | 1.0 |
Quickstart
One line, with uv installed. It fetches the vasuki program, downloads the model on the first run, and answers.
Mac with Apple silicon:
uvx --from git+https://github.com/rahul05ranjan/vasuki-vic vasuki ask "Total amount per vendor, highest first." -i invoices.csv --show-plan
macOS, Linux or Windows with Ollama:
uvx --from git+https://github.com/rahul05ranjan/vasuki-vic vasuki ask "Total amount per vendor, highest first." -i invoices.csv --show-plan --ollama
{
"status": "ok",
"columns": ["vendor", "total_amount"],
"row_count": 3,
"rows": [["Acme", 2500.0], ["Bolt", 500.5], ["Cora", 0]],
"plan": [
{"op": "aggregate", "group_by": ["vendor"], "metrics": [{"as": "total_amount", "fn": "sum", "column": "amount"}]},
{"op": "sort", "by": [{"column": "total_amount", "order": "desc"}]}
]
}
Intended use
- Answering questions about CSV, TSV, JSON and JSONL files whose answer is a table: totals, counts, averages, filters, rankings, joins of two tables, unit conversions.
- Handing tabular work to a small local model from an application or an agent, so the data stays on the machine and out of a larger model's context.
- Any setting where the answer must be auditable: the plan is short, typed, and says exactly which rows and columns were used.
Out of scope
- General chat, code generation, or free-text answers. The model replies with one JSON object and nothing else.
- Unattended use where a wrong table would cause harm. About one request in seven gets a plan that runs and returns the wrong table, with no warning.
- Work the plan format cannot express: SQL sources, time zones, currency conversion, charts.
Evaluation
Each task is scored by running the model's plan and comparing the resulting table with the expected one, so any plan that gives the right table counts. All figures are for this 6-bit model. The reference is gpt-oss-120b, a 120B-parameter model, given the full format description with every request.
| Test set | Tasks | VIC-2B | Reference model |
|---|---|---|---|
Questions real people wrote, from Spider's validation split (evals/real1) |
681 | 83.0% | 81.6% |
The same kind from a second source, BIRD's dev split (evals/real2) |
195 | 66.7% | 65.6% |
Hand-written requests that spell out their output (evals/test1) |
173 | 82.1% | 93.6% |
Development set (evals/v2) |
245 | 85.7% | 97.1% |
- The first two rows ignore column names, because those questions name none. On both, the two models cannot be told apart: 52 tasks are right only in VIC-2B against 43 only in the reference model on the first (p = 0.41), and 25 against 23 on the second (p = 0.89).
- VIC-2B was trained on questions from Spider's train split. No question or database of
evals/real1is among them, but the wording is familiar. The BIRD row is the check on that: other authors, other databases, nothing from BIRD trained on. - On requests that state exactly which columns to return and what to call them, the large model is ahead.
- VIC-2B reads about 685 tokens and writes about 77 per request on the BIRD set; the reference model reads 3,155 and writes 509.
- Speed: about 0.5 seconds per request on an M5 Pro once loaded.
The eval sets, the harness and the stored answers of both models are in the code repository, so every number can be recomputed.
Limitations and risks
- Wrong answers are silent. A wrong plan still runs and still returns a table. Read the plan before relying on an answer.
- What it gets wrong, read from its misses on 250 development questions of the real-people kind, most of them on two tables:
- after a join, grouping by a key and then asking for a column the grouping dropped, or grouping by a name where the question meant an id;
- a missing "distinct";
- a value the question spells differently from the data ("California" where the column holds
CA); - a filter left out, or applied to rows where the question meant a group total.
- Harder questions. Questions that lean on what a column means, as BIRD's do, are right two times in three.
- Quantisation. Use 6 bits or more. On an earlier checkpoint, 4 bits cost almost 6 points.
- Reasoning must be off. The model was trained to reply with the plan alone. Asked to reason first, it argues with the request and returns no plan. Pass
enable_thinking=False; vasuki does this for you.
Use from Python
import json
from mlx_lm import load, generate
from vasuki import load_table, run_plan
from vasuki.request import VIC_SYSTEM, planner_input
model, tokenizer = load("rahul05ranjan/VIC-2B")
tables = {"orders": load_table("orders.csv")}
request = "Total items per channel, as total_items."
messages = [
{"role": "system", "content": VIC_SYSTEM},
{"role": "user", "content": json.dumps(planner_input(request, tables), separators=(",", ":"))},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, enable_thinking=False)
reply = json.loads(generate(model, tokenizer, prompt=prompt, max_tokens=400))
if reply["status"] == "plan":
print(run_plan(reply, tables).report())
else:
print(reply) # a clarifying question, or a statement that the request is out of scope
The reply is one JSON object: a plan, a needs_clarification question, or an unsupported reason.
Input and output
Input. A system prompt (vasuki.request.VIC_SYSTEM) and one user message: a JSON object holding the request and, for each table, its row count, column names and types, null counts, three sample rows, and the distinct values of text columns that have twenty or fewer.
Output. One JSON object:
{"status": "plan", "steps": [{"op": "filter", "where": "amount > 100"}]}
{"status": "needs_clarification", "question": "Is GST included in amount?"}
{"status": "unsupported", "reason": "needs a currency exchange rate"}
The plan format has 12 operations and is described in the code repository.
Training
- Method. LoRA, rank 64, on the attention and MLP projections of every layer of
Qwen/Qwen3.5-2B(text part, half precision), with the prompt masked out of the loss. The adapter was merged and the result quantised to 6 bits for MLX. - Stage one. 1.09 passes over 17,323 synthetic examples, about seven hours on one T4 GPU.
- Stage two. Continued at half the learning rate for 10,077 more examples, 4.5 hours on one T4 GPU.
- Synthetic data, machine-verified. One model invents a small dataset, a second writes requests and plans on it, and a third solves each request in pandas from the request alone. An example is kept only when both produce the same table. Requests were then rewritten as short statements of intent and in everyday wording, and verified again each time.
- Real questions. 4,046 questions from the train split of Spider (Yu et al., 2018, CC BY-SA 4.0), each with a plan written by a large model and kept only when the plan's result equals the result of Spider's own SQL, mixed with 30% of the synthetic data.
Licence
Released under the Apache License 2.0, the licence of the base model. Part of the training data is derived from Spider (Yu et al., 2018), which is CC BY-SA 4.0; the eval sets built from Spider and from BIRD (Li et al., 2023) are distributed under that licence in the code repository. The terms of the models that wrote the synthetic training data, and of the service they ran on, were read and nothing was found against training on their output. This is a reading of those documents, not legal advice.
Citation
@software{vasuki_vic_2026,
title = {vasuki and VIC-2B: typed plans, a deterministic engine and a small planner model for tabular questions},
author = {rahul05ranjan},
year = {2026},
url = {https://github.com/rahul05ranjan/vasuki-vic}
}
Contact
Questions and bug reports: GitHub issues. Security reports: see the repository's security policy.
- Downloads last month
- 3
6-bit
Model tree for rahul05ranjan/VIC-2B
Dataset used to train rahul05ranjan/VIC-2B
Evaluation results
- Execution accuracy, column names ignored on Spider validation, vasuki real1 (681 questions)validation set self-reported83.000
- Execution accuracy, column names ignored on BIRD dev, vasuki real2 (195 questions)self-reported66.700
- Execution accuracy on vasuki test1 (173 hand-written requests)self-reported82.100