Instructions to use jgeiser/account-intelligence-7b-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ELM
How to use jgeiser/account-intelligence-7b-v1 with ELM:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jgeiser/account-intelligence-7b-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jgeiser/account-intelligence-7b-v1:Q4_K_M # Run inference directly in the terminal: llama cli -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jgeiser/account-intelligence-7b-v1:Q4_K_M # Run inference directly in the terminal: llama cli -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jgeiser/account-intelligence-7b-v1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jgeiser/account-intelligence-7b-v1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Use Docker
docker model run hf.co/jgeiser/account-intelligence-7b-v1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use jgeiser/account-intelligence-7b-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jgeiser/account-intelligence-7b-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgeiser/account-intelligence-7b-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jgeiser/account-intelligence-7b-v1:Q4_K_M
- Ollama
How to use jgeiser/account-intelligence-7b-v1 with Ollama:
ollama run hf.co/jgeiser/account-intelligence-7b-v1:Q4_K_M
- Unsloth Desktop
- Pi
How to use jgeiser/account-intelligence-7b-v1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jgeiser/account-intelligence-7b-v1:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jgeiser/account-intelligence-7b-v1 with Docker Model Runner:
docker model run hf.co/jgeiser/account-intelligence-7b-v1:Q4_K_M
- Lemonade
How to use jgeiser/account-intelligence-7b-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jgeiser/account-intelligence-7b-v1:Q4_K_M
Run and chat with the model
lemonade run user.account-intelligence-7b-v1-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use jgeiser/account-intelligence-7b-v1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jgeiser/account-intelligence-7b-v1:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jgeiser/account-intelligence-7b-v1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jgeiser/account-intelligence-7b-v1:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jgeiser/account-intelligence-7b-v1:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
account-intelligence-7b-v1
Jeff Geiser, Zenlayer
A fine-tuned Expert Language Model (ELM) that synthesizes structured account intelligence briefs from raw enterprise data sources. Given a bundle of account signals โ support tickets, usage metrics, contract terms, stakeholder notes โ it produces a complete, source-attributed JSON brief covering health, risk, expansion signals, and QBR readiness across six surfaces.
This model is a demonstration of the ELM methodology: small, task-specific language models fine-tuned to emit structured JSON rather than prose, constrained at inference time to guarantee schema validity.
Model Details
| Base model | Qwen/Qwen2.5-7B-Instruct |
| Fine-tuning method | LoRA (rank 16, ฮฑ 16) via Unsloth + TRL SFTTrainer |
| Training rounds | 10 |
| Training examples | 1,127 Path B (source bundle โ structured brief) pairs |
| Max sequence length | 16,384 tokens |
| Output format | Constrained JSON (Outlines / llama.cpp grammar) |
| Releases | LoRA adapter (fp16) + Q4_K_M GGUF |
Evaluation Results
Evaluated on 60 held-out examples never seen during training.
| Model | M1 Schema Adherence | M3 Section Coverage | M5 Source Attribution |
|---|---|---|---|
| Base Qwen2.5-7B-Instruct (untuned) | 100% | 28.1% | 69.2% |
| Round 10 LoRA adapter (fp16, Outlines) | 100% | 95.8% | 92.7% |
| Round 10 Q4_K_M GGUF (llama.cpp grammar) | 100% | 99.2% | 96.1% |
M1 โ Schema Adherence: Output parses as valid JSON matching the required schema. 100% for all rows because constrained decoding enforces schema validity at the token level โ it is a floor, not a differentiator.
M3 โ Section Coverage: Of all sections present in the gold-standard brief, how many did the model fill with substantive content? The base model writes skeletal 2-3KB briefs; the fine-tuned model writes complete 14-16KB briefs. The 28% โ 99% gain is the core fine-tuning result.
M5 โ Source Attribution: Of the sources provided in the input bundle, how many are correctly cited in the output? Fine-tuning on source-attributed training pairs pushed this from 69% to 96%.
Note: M2 (hallucination rate) has not yet been measured. M1 being 100% across all rows reflects constrained decoding, not model capability.
Intended Use
This model is designed for enterprise account teams who need structured intelligence briefs synthesized from multiple data sources. It covers six surfaces: account health, escalation risk, expansion signals, stakeholder map, QBR readiness, and competitive exposure.
It is not a general-purpose model. It is an ELM โ trained to do one thing well, emit structured JSON that downstream systems can consume without parsing prose.
Suitable for:
- Automated brief generation before QBRs or executive reviews
- Integration into account intelligence pipelines
- Demonstration of the ELM fine-tuning methodology
Not suitable for:
- General conversation or instruction following
- Tasks outside account intelligence synthesis
- Domains outside enterprise telco / infrastructure accounts without fine-tuning on domain-specific data
Usage
GGUF (recommended for local inference)
llama-cli \
-m account-intelligence-7b-v1-Q4_K_M.gguf \
--json-schema-file schema.json \
-p "<your prompt>" \
-n 8192 \
--temp 0.1
The GGUF runs on CPU or GPU. Apple Silicon supported via Metal (-ngl 99 to offload all layers).
LoRA Adapter (fp16, requires GPU)
from unsloth import FastLanguageModel
from peft import PeftModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="Qwen/Qwen2.5-7B-Instruct",
max_seq_length=16384,
load_in_4bit=True,
)
model = PeftModel.from_pretrained(model, "jgeiser/account-intelligence-7b-v1")
FastLanguageModel.for_inference(model)
Constrained decoding via Outlines is strongly recommended. Without it, schema adherence degrades significantly.
Training Methodology
Trained using the ELM (Expert Language Model) methodology โ small, specialized models that emit structured output for specific enterprise tasks.
- Path B training data: (source bundle โ structured brief) pairs. The model learns to synthesize, not summarize.
- Completion-only loss masking: Loss computed only on JSON output tokens, not the prompt.
- Constrained decoding: Outlines (training/eval) and llama.cpp JSON schema grammar (GGUF) enforce schema validity via FSM-guided logit masking.
- Sequence length: MAX_SEQ_LENGTH=16384. An earlier run at 8192 silently truncated 27% of training examples.
10 training rounds. Rounds 1-5 used incorrect data format. Rounds 6-10 used correct Path B format. Round 10 with Outlines: M1=100%, M3=95.8%, M5=92.7%.
Training Data
Proprietary internal account data, not released. 1,127 Path B examples: structured source bundles paired with human-verified account intelligence briefs across enterprise telco and infrastructure accounts. 60 held-out examples used for all reported metrics.
Limitations
- Domain-specific: trained on enterprise telco/infrastructure accounts; performance on other verticals is untested
- M2 (hallucination rate) not yet measured
- Requires constrained decoding for reliable schema adherence
- Q4_K_M quantization outperformed fp16 adapter on M3/M5 in eval; may reflect eval variance
Files
| File | Size | Description |
|---|---|---|
adapter_model.safetensors |
~100MB | LoRA adapter weights (Round 10) |
adapter_config.json |
โ | LoRA configuration |
account-intelligence-7b-v1-Q4_K_M.gguf |
4.68GB | Q4_K_M quantized GGUF for llama.cpp / Ollama |
Part of ongoing ELM methodology work. Built at Zenlayer.
- Downloads last month
- 10
4-bit