account-intelligence-7b-v1

Jeff Geiser, Zenlayer

A fine-tuned Expert Language Model (ELM) that synthesizes structured account intelligence briefs from raw enterprise data sources. Given a bundle of account signals โ€” support tickets, usage metrics, contract terms, stakeholder notes โ€” it produces a complete, source-attributed JSON brief covering health, risk, expansion signals, and QBR readiness across six surfaces.

This model is a demonstration of the ELM methodology: small, task-specific language models fine-tuned to emit structured JSON rather than prose, constrained at inference time to guarantee schema validity.


Model Details

Base model Qwen/Qwen2.5-7B-Instruct
Fine-tuning method LoRA (rank 16, ฮฑ 16) via Unsloth + TRL SFTTrainer
Training rounds 10
Training examples 1,127 Path B (source bundle โ†’ structured brief) pairs
Max sequence length 16,384 tokens
Output format Constrained JSON (Outlines / llama.cpp grammar)
Releases LoRA adapter (fp16) + Q4_K_M GGUF

Evaluation Results

Evaluated on 60 held-out examples never seen during training.

Model M1 Schema Adherence M3 Section Coverage M5 Source Attribution
Base Qwen2.5-7B-Instruct (untuned) 100% 28.1% 69.2%
Round 10 LoRA adapter (fp16, Outlines) 100% 95.8% 92.7%
Round 10 Q4_K_M GGUF (llama.cpp grammar) 100% 99.2% 96.1%

M1 โ€” Schema Adherence: Output parses as valid JSON matching the required schema. 100% for all rows because constrained decoding enforces schema validity at the token level โ€” it is a floor, not a differentiator.

M3 โ€” Section Coverage: Of all sections present in the gold-standard brief, how many did the model fill with substantive content? The base model writes skeletal 2-3KB briefs; the fine-tuned model writes complete 14-16KB briefs. The 28% โ†’ 99% gain is the core fine-tuning result.

M5 โ€” Source Attribution: Of the sources provided in the input bundle, how many are correctly cited in the output? Fine-tuning on source-attributed training pairs pushed this from 69% to 96%.

Note: M2 (hallucination rate) has not yet been measured. M1 being 100% across all rows reflects constrained decoding, not model capability.


Intended Use

This model is designed for enterprise account teams who need structured intelligence briefs synthesized from multiple data sources. It covers six surfaces: account health, escalation risk, expansion signals, stakeholder map, QBR readiness, and competitive exposure.

It is not a general-purpose model. It is an ELM โ€” trained to do one thing well, emit structured JSON that downstream systems can consume without parsing prose.

Suitable for:

  • Automated brief generation before QBRs or executive reviews
  • Integration into account intelligence pipelines
  • Demonstration of the ELM fine-tuning methodology

Not suitable for:

  • General conversation or instruction following
  • Tasks outside account intelligence synthesis
  • Domains outside enterprise telco / infrastructure accounts without fine-tuning on domain-specific data

Usage

GGUF (recommended for local inference)

llama-cli \
  -m account-intelligence-7b-v1-Q4_K_M.gguf \
  --json-schema-file schema.json \
  -p "<your prompt>" \
  -n 8192 \
  --temp 0.1

The GGUF runs on CPU or GPU. Apple Silicon supported via Metal (-ngl 99 to offload all layers).

LoRA Adapter (fp16, requires GPU)

from unsloth import FastLanguageModel
from peft import PeftModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="Qwen/Qwen2.5-7B-Instruct",
    max_seq_length=16384,
    load_in_4bit=True,
)
model = PeftModel.from_pretrained(model, "jgeiser/account-intelligence-7b-v1")
FastLanguageModel.for_inference(model)

Constrained decoding via Outlines is strongly recommended. Without it, schema adherence degrades significantly.


Training Methodology

Trained using the ELM (Expert Language Model) methodology โ€” small, specialized models that emit structured output for specific enterprise tasks.

  • Path B training data: (source bundle โ†’ structured brief) pairs. The model learns to synthesize, not summarize.
  • Completion-only loss masking: Loss computed only on JSON output tokens, not the prompt.
  • Constrained decoding: Outlines (training/eval) and llama.cpp JSON schema grammar (GGUF) enforce schema validity via FSM-guided logit masking.
  • Sequence length: MAX_SEQ_LENGTH=16384. An earlier run at 8192 silently truncated 27% of training examples.

10 training rounds. Rounds 1-5 used incorrect data format. Rounds 6-10 used correct Path B format. Round 10 with Outlines: M1=100%, M3=95.8%, M5=92.7%.


Training Data

Proprietary internal account data, not released. 1,127 Path B examples: structured source bundles paired with human-verified account intelligence briefs across enterprise telco and infrastructure accounts. 60 held-out examples used for all reported metrics.


Limitations

  • Domain-specific: trained on enterprise telco/infrastructure accounts; performance on other verticals is untested
  • M2 (hallucination rate) not yet measured
  • Requires constrained decoding for reliable schema adherence
  • Q4_K_M quantization outperformed fp16 adapter on M3/M5 in eval; may reflect eval variance

Files

File Size Description
adapter_model.safetensors ~100MB LoRA adapter weights (Round 10)
adapter_config.json โ€” LoRA configuration
account-intelligence-7b-v1-Q4_K_M.gguf 4.68GB Q4_K_M quantized GGUF for llama.cpp / Ollama

Part of ongoing ELM methodology work. Built at Zenlayer.

Downloads last month
10
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jgeiser/account-intelligence-7b-v1

Base model

Qwen/Qwen2.5-7B
Adapter
(2813)
this model