vasuki: ask a table a question in plain words

VIC-2B

A 1.55 GB model that turns a question about a data file into a typed plan, which a deterministic engine then runs on your machine.

VIC (Vasuki Indicus Compute) is the planner model of vasuki. Given a request in plain words and a short profile of the tables it refers to, VIC-2B writes a plan of typed steps: filter, derive, aggregate, join, convert units and so on. The vasuki engine checks the plan against the columns and runs it. The model never computes a result, never writes code, and never sees the whole file.

Developer rahul05ranjan
Model type Causal language model, fine-tuned to write vasuki plans as JSON
Base model Qwen/Qwen3.5-2B, text part only (1.9B parameters)
This repository MLX, 6-bit, 1.55 GB, for Apple silicon
Other formats VIC-2B-GGUF: 8-bit GGUF, 2.0 GB, for Ollama and llama.cpp on any platform
Languages English; Hinglish to a lesser degree
Licence Apache 2.0
Code github.com/rahul05ranjan/vasuki-vic
Version 1.0

Quickstart

One line, with uv installed. It fetches the vasuki program, downloads the model on the first run, and answers.

Mac with Apple silicon:

uvx --from git+https://github.com/rahul05ranjan/vasuki-vic vasuki ask "Total amount per vendor, highest first." -i invoices.csv --show-plan

macOS, Linux or Windows with Ollama:

uvx --from git+https://github.com/rahul05ranjan/vasuki-vic vasuki ask "Total amount per vendor, highest first." -i invoices.csv --show-plan --ollama
{
  "status": "ok",
  "columns": ["vendor", "total_amount"],
  "row_count": 3,
  "rows": [["Acme", 2500.0], ["Bolt", 500.5], ["Cora", 0]],
  "plan": [
    {"op": "aggregate", "group_by": ["vendor"], "metrics": [{"as": "total_amount", "fn": "sum", "column": "amount"}]},
    {"op": "sort", "by": [{"column": "total_amount", "order": "desc"}]}
  ]
}

Intended use

  • Answering questions about CSV, TSV, JSON and JSONL files whose answer is a table: totals, counts, averages, filters, rankings, joins of two tables, unit conversions.
  • Handing tabular work to a small local model from an application or an agent, so the data stays on the machine and out of a larger model's context.
  • Any setting where the answer must be auditable: the plan is short, typed, and says exactly which rows and columns were used.

Out of scope

  • General chat, code generation, or free-text answers. The model replies with one JSON object and nothing else.
  • Unattended use where a wrong table would cause harm. About one request in seven gets a plan that runs and returns the wrong table, with no warning.
  • Work the plan format cannot express: SQL sources, time zones, currency conversion, charts.

Evaluation

Each task is scored by running the model's plan and comparing the resulting table with the expected one, so any plan that gives the right table counts. All figures are for this 6-bit model. The reference is gpt-oss-120b, a 120B-parameter model, given the full format description with every request.

Test set Tasks VIC-2B Reference model
Questions real people wrote, from Spider's validation split (evals/real1) 681 83.0% 81.6%
The same kind from a second source, BIRD's dev split (evals/real2) 195 66.7% 65.6%
Hand-written requests that spell out their output (evals/test1) 173 82.1% 93.6%
Development set (evals/v2) 245 85.7% 97.1%
  • The first two rows ignore column names, because those questions name none. On both, the two models cannot be told apart: 52 tasks are right only in VIC-2B against 43 only in the reference model on the first (p = 0.41), and 25 against 23 on the second (p = 0.89).
  • VIC-2B was trained on questions from Spider's train split. No question or database of evals/real1 is among them, but the wording is familiar. The BIRD row is the check on that: other authors, other databases, nothing from BIRD trained on.
  • On requests that state exactly which columns to return and what to call them, the large model is ahead.
  • VIC-2B reads about 685 tokens and writes about 77 per request on the BIRD set; the reference model reads 3,155 and writes 509.
  • Speed: about 0.5 seconds per request on an M5 Pro once loaded.

The eval sets, the harness and the stored answers of both models are in the code repository, so every number can be recomputed.

Limitations and risks

  • Wrong answers are silent. A wrong plan still runs and still returns a table. Read the plan before relying on an answer.
  • What it gets wrong, read from its misses on 250 development questions of the real-people kind, most of them on two tables:
    • after a join, grouping by a key and then asking for a column the grouping dropped, or grouping by a name where the question meant an id;
    • a missing "distinct";
    • a value the question spells differently from the data ("California" where the column holds CA);
    • a filter left out, or applied to rows where the question meant a group total.
  • Harder questions. Questions that lean on what a column means, as BIRD's do, are right two times in three.
  • Quantisation. Use 6 bits or more. On an earlier checkpoint, 4 bits cost almost 6 points.
  • Reasoning must be off. The model was trained to reply with the plan alone. Asked to reason first, it argues with the request and returns no plan. Pass enable_thinking=False; vasuki does this for you.

Use from Python

import json
from mlx_lm import load, generate
from vasuki import load_table, run_plan
from vasuki.request import VIC_SYSTEM, planner_input

model, tokenizer = load("rahul05ranjan/VIC-2B")
tables = {"orders": load_table("orders.csv")}
request = "Total items per channel, as total_items."

messages = [
    {"role": "system", "content": VIC_SYSTEM},
    {"role": "user", "content": json.dumps(planner_input(request, tables), separators=(",", ":"))},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, enable_thinking=False)
reply = json.loads(generate(model, tokenizer, prompt=prompt, max_tokens=400))

if reply["status"] == "plan":
    print(run_plan(reply, tables).report())
else:
    print(reply)  # a clarifying question, or a statement that the request is out of scope

The reply is one JSON object: a plan, a needs_clarification question, or an unsupported reason.

Input and output

Input. A system prompt (vasuki.request.VIC_SYSTEM) and one user message: a JSON object holding the request and, for each table, its row count, column names and types, null counts, three sample rows, and the distinct values of text columns that have twenty or fewer.

Output. One JSON object:

{"status": "plan", "steps": [{"op": "filter", "where": "amount > 100"}]}
{"status": "needs_clarification", "question": "Is GST included in amount?"}
{"status": "unsupported", "reason": "needs a currency exchange rate"}

The plan format has 12 operations and is described in the code repository.

Training

  • Method. LoRA, rank 64, on the attention and MLP projections of every layer of Qwen/Qwen3.5-2B (text part, half precision), with the prompt masked out of the loss. The adapter was merged and the result quantised to 6 bits for MLX.
  • Stage one. 1.09 passes over 17,323 synthetic examples, about seven hours on one T4 GPU.
  • Stage two. Continued at half the learning rate for 10,077 more examples, 4.5 hours on one T4 GPU.
  • Synthetic data, machine-verified. One model invents a small dataset, a second writes requests and plans on it, and a third solves each request in pandas from the request alone. An example is kept only when both produce the same table. Requests were then rewritten as short statements of intent and in everyday wording, and verified again each time.
  • Real questions. 4,046 questions from the train split of Spider (Yu et al., 2018, CC BY-SA 4.0), each with a plan written by a large model and kept only when the plan's result equals the result of Spider's own SQL, mixed with 30% of the synthetic data.

Licence

Released under the Apache License 2.0, the licence of the base model. Part of the training data is derived from Spider (Yu et al., 2018), which is CC BY-SA 4.0; the eval sets built from Spider and from BIRD (Li et al., 2023) are distributed under that licence in the code repository. The terms of the models that wrote the synthetic training data, and of the service they ran on, were read and nothing was found against training on their output. This is a reading of those documents, not legal advice.

Citation

@software{vasuki_vic_2026,
  title  = {vasuki and VIC-2B: typed plans, a deterministic engine and a small planner model for tabular questions},
  author = {rahul05ranjan},
  year   = {2026},
  url    = {https://github.com/rahul05ranjan/vasuki-vic}
}

Contact

Questions and bug reports: GitHub issues. Security reports: see the repository's security policy.

Downloads last month
3
Safetensors
Model size
2B params
Tensor type
U32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rahul05ranjan/VIC-2B

Finetuned
Qwen/Qwen3.5-2B
Quantized
(240)
this model
Quantizations
1 model

Dataset used to train rahul05ranjan/VIC-2B

Evaluation results

  • Execution accuracy, column names ignored on Spider validation, vasuki real1 (681 questions)
    validation set self-reported
    83.000
  • Execution accuracy, column names ignored on BIRD dev, vasuki real2 (195 questions)
    self-reported
    66.700
  • Execution accuracy on vasuki test1 (173 hand-written requests)
    self-reported
    82.100