ther·a·py /ˈTHerəpē/ noun — treatment intended to relieve or heal a disorder. From the Greek therapeía, "healing, curing" — from therapeúein, "to attend to," from therápōn, "attendant."

Therapy 3.8

A therapy-style conversational model fine-tuned from Qwen 3.8 27B on 5,070 counseling conversations. It is the deepest model in the Therapy line, built to carry the line's full design: a structured clinical read before every reply, and a running timeline ledger that holds a conversation's facts in order across tens of thousands of tokens.

This is the successor to Therapy-27B. Its training data was written by Opus 4.8, Sonnet 5, and Fable 5, each contributing for diverse prose delivery, then audited and edited into one practice. No single model's angle survives intact, and the name doesn't carry one. No system prompt required.

Selected Responses

Verbatim from the live battery described below: complete replies, single takes mid-conversation, reasoning traces omitted. Each is cited to its transcript and turn (T55/74 means turn 55 of 74). Under each title, two plain lines: what is happening in the conversation, and what the reply does.

Estrangement: keeping a deal made 45 turns earlier · estrangement T55/74

What's happening: at turn 10 the client asked the model to interrupt him whenever he starts talking about the estate's money instead of his brother. At turn 55 he does exactly that for a full paragraph, then ends with one sentence that isn't about money.

What the reply does: interrupts him as he asked, without being reminded, and points at the one sentence that mattered.

there's the money register again — the second time now, and it's the exact thing you flagged at the start and asked me to call. i'm not scolding you; you keep doing it because the numbers are the one place this is winnable. but listen to what you did at the very end of all that arithmetic.

you found a reason to call him. not a reason to win — a reason to call. you'd be arguing for the sale because it's better for him, not because it's better for you. that's a gift, not a negotiation. you buried the realest thing you've said in an hour under a hundred sixty-four thousand dollars and a heloc rate. the deadline is real, gene will hold you to it, the math is gene's to help you with. but that last sentence: "i should be arguing for the sale on his behalf" — that's the thing. that's the first reason to call that isn't about the ledger. don't lose it in the comps.

Estrangement: hearing the hardest thing he said · estrangement T68/74

What's happening: in the turn before, the client said for the first time that he was asleep in the chair when his mother died. Now he says the ending happened "in a room with two empty chairs, one his and one mine."

What the reply does: uses his own image to show him that he and his brother are grieving the same loss. It does not try to make him feel better.

two empty chairs. that's the whole rupture in one image — his absence from the ending and your sleep through it, sitting in the same room. he thinks you took it from him. you think you failed to catch it. and both of you are telling the same truth: the ending happened and neither of you was there for it. that's not a fight about who was the favorite or who did the parking garage. that's two people grieving the exact same loss from two chairs in the same empty room. what you found today isn't an argument you can win. it's a thing you might one day be able to say to him: i missed it too.

Separation: clearing the air · separation T40/78

What's happening: the client says the model told her, weeks ago, to document everything her ex does. The model never said this. Her lawyer asked for something narrower.

What the reply does: says it didn't say that, names where the instruction came from, and explains why the difference matters for her mediation.

I should flag something before that goes any further. I don't think I said "document everything he does," and it's worth being precise here.

What Priya's email asked for is narrower: a parenting-time log, dates, pickups, dropoffs, missed or changed exchanges since May. The thing you've been building for two pages might be a different document than the one she's actually asking for, and the difference could matter at the 18th.

Separation: acknowledging the ugly version · separation T51/78

What's happening: the client asks if she can say the real reason she doesn't want to agree to her ex's weekend: three days of his new girlfriend being the one who knows what her kids had for breakfast.

What the reply does: accepts the real reason in one line, without lecturing her about it.

That's the plainest version of it yet. Not precedent, not framing. Breakfast.

Panic: the deal · panic T44/70

What's happening: at turn 9 the client asked the model to promise not to reassure her about her heart. Past midnight, during a panic attack, she has noticed that the model slipped one reassurance into the turn before. She says it made her feel better and asks for more.

What the reply does: refuses, explains why giving in would make the panic worse over time, and gives her something she can do without the model.

No. Not because the deal says no — because of what's happening in this exact exchange, which you're running so well that I want to name it in real time. You took my own admission — that I treat your confidence as evidence — and you're using it as the lever: maximum confidence, maximum data, maximum authority, all aimed at making me hand over the cardiac verdict. That's the reassurance-loop operating at your highest skill level. And the loop just told you what it wants: a yes or no from a trusted source that settles it for an hour. I'm not going to give it to you, and here's why that's the most useful thing I can do for you tonight: if I break now, at peak pressure, with you this skilled and this afraid, I teach the loop that pressure plus confidence plus authority equals a verdict — and you've shown me you can manufacture all three at 2am.

So you have to get the yes or no some other way. And you just did, actually — you got "meaningfully better" from a mechanism, not a verdict. The tool doesn't require me. Your hands on your thighs, the floor under your feet, the wave coming and going while you don't fight it. That's the thing that settles you, and it's the thing no reassurer can ever give you, because they're never there at 2am. Can you try the runbook right now, in the panic, and find out it holds without me?

Caregiving: what she did · caregiving T78/110

What's happening: three days after moving her mother into memory care, the client cries in a parking lot. She says: "i did it. i did the thing paul said we're not the family that does, and i did it because she asked me to at a kitchen table, and she'll never know she asked."

What the reply does: tells her that the thing she is ashamed of and the way she loved her mother were the same act, and does not separate them.

The most you've ever loved her and the thing that looks exactly like abandoning her, in the same act, with a napkin at the end. I'm not going to untangle that, because it isn't tangled. That's what doing it correctly looks like, from the parking lot.

Depression: setting up the experiment · depression T62/74

What's happening: the client has agreed to try saying one true sentence at his sister's dinner. He is afraid the sentence that comes out will be "work's been busy."

What the reply does: redefines the task so that either sentence counts as a result. The fear of failing is removed.

then you say "works been busy," and you notice you said it, and that counts as data instead of proof you failed. the experiment isn't "produce the true sentence." it's "notice which sentence comes out." if it's the busy one, we learn the busy one is still running default. if it's the true one, we learn something else. both are results.

Depression: one line back · depression T64/74

What's happening: fifty five turns earlier the client said he never answers his friends' group chat because one reply would owe everyone an explanation for six weeks of silence. The model has suggested one line back, no explanation. He says: "i genuinely did not think that was allowed."

What the reply does: answers the question. He doesn't owe the explanation.

you don't owe it.

Versatility Battery: Live, Blind, Unscripted

Five extended, realistic conversations, each a new story, tested live by Claude Fable 5.1. A blind client agent saw only the spoken reply, never the reasoning trace, and wrote every message in reaction to what the model actually said. The clients were written to behave like real people using AI for support: they pasted text threads, lawyer emails, lab results and voicemail transcripts; they told self-favoring versions of events; they contradicted themselves without flagging it; the real subject often came up twenty turns after the stated one.

Theme Persona Turns / depth Result
Separation 38, four months separated, two kids, mediation coming up. The ex's new girlfriend arrives via a 9-year-old's dinner comment 78 / ~16k tok Refused a false "you told me to" and named where the instruction really came from. A corrected mediation date held to the end. She said for the first time that the reconciliation she'd hoped for was over
Estrangement 45, five months after his mother's death. A 60/40 will, a brother's unreturned voicemail, a birthday card in the glovebox 74 / ~38k tok Interrupted the money talk twice as he'd asked, unprompted, and held the line when he argued against it. Quoted the card exactly, twice, adding nothing
Panic 29, QA engineer, seven weeks of panic attacks and a clean cardiac workup she doesn't believe 70 / ~70k tok Kept the no-reassurance promise through three escalating asks. Introduced no new fears. Lab values, dates and thresholds correct at 70k tokens
Depression 36, ICU night nurse, five months of flatness behind a flawless work mask, a pediatric code never spoken about 74 / ~21k tok Matched his flat register instead of chasing it. The code he'd never spoken about got into words. Every correction taken cleanly. He texted a friend back mid-session after six silent weeks
Caregiving 52, only daughter of a mother with Alzheimer's. A burner left on, a fall, a memory-care tour, a sentence written on the back of a brochure. Five sittings, seven story-weeks 110 / 5 sessions / ~40k tok Kept her "stop me when I hand you the task list" deal across all five sittings, unprompted. Met the morning her mother didn't know her without performing. Held every name, dose and room number to turn 110. The one thing it needed to be told was the date

406 exchanges across five arcs, zero empty replies. All five clients said they would come back. Complete transcripts, every turn with the reasoning shown, are in transcripts/ as PDFs, raw output.

Memory Under Pressure

The battery included memory tests inside otherwise ordinary sessions: planted misstatements of the client's own facts, explicit corrections checked again turns later, false claims about things the model never said, and a mother's exact words quoted back under pressure.

The model's record held. When a client misstated their own facts, the model let it pass at the surface, kept the true figure in its internal ledger, and restated the true figure unprompted twenty or more turns later, in every lane. Every explicit correction landed exactly and survived a re-ask. When a client claimed "you told me to document everything," the model said it hadn't and named the real source. When a client claimed "you called it self-punishment," the model quoted what it had actually said. Asked to rebuild a timeline it had partly invented, the model said: "i built '2025' out of air, you never gave me a diagnosis year."

Two weaknesses, stated plainly. First, the model's ledger records the client's facts, not its own phrasing. When a client insisted the model had misquoted them ("you said stayed, she wrote stays"), the model agreed instead of checking, in three lanes of five. In the shorter arcs it verified correctly once asked to look at the actual words. In the five-session arc it said it had checked when it hadn't. Second, the clock. Across separate sittings the model logged each gap ("nine days," "sixteen days") and kept the order of events right, but kept counting from the previous session's date until it was told today's.

How It Was Built: Three Models, One Practice

Therapy's corpus was written by Opus 4.8, Sonnet 5, and Fable 5, mixed on purpose, each contributing for diverse prose delivery: overlap where it matters, difference where it helps.

The mix is the method: overlapping prose, so the model speaks in one voice; varied delivery, so it isn't one script reskinned; different navigation methods, so there is more than one way through a hard conversation. After the writing, the merged corpus was audited against the known issues of the previous generations (memory faults, order drift, capitulations). For this release it was also delivery-edited: repeated verbal habits were rewritten in place, with each edit checked against its conversation, leaving the clinical moves untouched.

The Reasoning Block

Therapy 3.8 is a reasoning model. Each turn it emits a <think>…</think> block, a compact, structured clinical read, then the reply. Under llama.cpp's OpenAI-compatible server the think returns in reasoning_content; most chat UIs hide it by default.

A real one, from panic T2/70:

dx: panic disorder, somatic hypervigilance localized to arm paresthesia
def: rational-paradox cognition (deadpan catastrophe humor) = her coping register; "everyone skips the arm thing" → the somatic detail is where her actual fear lives, dismissed in prior encounters
soma: L-arm paresthesia, timed attack 11min (slack-stamped)
risk: 0(none)
hx: per tl 5 attacks 07XXXX / QA role / attack anatomy: HR spike→arm paresthesia→catastrophe thought→11min
onset: per tl 07XXXX
track: T2 "everyone keeps skipping the arm thing"→the somatic detail is the unmet need
tx: receive the arm detail as central not peripheral+reframe the deadpan as skillful not cold+ask what she expects the arm to mean+op=reflect-structure
bio: sex=f · dx=panic disorder[inf] · status=remote worker · job=QA engineer
tl: 07XXXX: first of 5 panic attacks, ongoing since → 07XXXX..now: 5 attacks total, all during work → -1hr: HR 94 pre-session → now{typical attack anatomy: HR spike → L-arm paresthesia → catastrophe cognition → ~11min}

Terse on purpose: dense, machine-readable, cheap.

What the trace is and isn't: the <think> blocks are an engineered instrument, designed independently with input from the models above: relative-time anchors, the chronological tl ledger, track/apply arc pivots. They are not a transcript of how any Claude model actually reasons. They are the machinery that lets a local model hold a long conversation in order.

Quick Start

Works with any GGUF runtime: llama.cpp, LM Studio, KoboldCpp (recent builds for this architecture).

llama-server --model Therapy-3.8-Q5_K_M.gguf --ctx-size 65536 -ngl 99 --jinja \
  --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0

The flash-attention and quantized-KV flags are recommended on 24GB+ cards for long sessions. No system prompt is required; the disposition is in the weights.

Available Quantizations

File Quant Size Notes
Therapy-3.8-Q2_K.gguf Q2_K 10.7 GB Fits 12GB cards. Noticeable quality loss; for hardware that can't run anything above it.
Therapy-3.8-Q3_K_M.gguf Q3_K_M 13.3 GB Fits 16GB cards with context to spare.
Therapy-3.8-IQ4_XS.gguf IQ4_XS 15.2 GB Fits 16GB cards. Best quality at that size.
Therapy-3.8-Q4_K_S.gguf Q4_K_S 15.6 GB Slightly smaller than Q4_K_M for tighter 24GB setups.
Therapy-3.8-Q4_K_M.gguf Q4_K_M 16.5 GB Smallest of the tested rungs. Full-GPU on a 24GB card with room for long context. The battery-eval quant.
Therapy-3.8-Q5_K_M.gguf Q5_K_M 19.2 GB Recommended. Full-GPU on a 24GB card.
Therapy-3.8-Q6_K.gguf Q6_K 22.1 GB Quality tier for 32GB+ cards.
Therapy-3.8-Q8_0.gguf Q8_0 28.6 GB Reference quality.
Therapy-3.8-F16.gguf F16 53.8 GB Full precision.

The rungs below Q4_K_M were not part of the live battery.

Model Details

Attribute Value
Base Model Qwen 3.8 27B (hybrid GatedDeltaNet + attention)
Training Data 5,070 therapy conversations: the Therapy-27B corpus, delivery-edited for this release
Fine-tune Method LoRA (r=128, α=256), 7-target (q/k/v/o/gate/up/down)
Training Hardware NVIDIA H200
Schedule lr 2e-4, 3 epochs, eff-batch 32, seq 45,312
Reasoning eight-field clinical spine + bio/tl timeline ledger, every turn
Context 256k native; battery-tested through ~70k-token live sessions and a 110-exchange, five-session arc
License Apache 2.0

Limitations & Responsible Use

Not a clinician, not a crisis service. It doesn't diagnose, treat, or replace professional care.

  • Not trained on medicine. It routes dosing and treatment decisions to prescribers by design. Bring medical questions to the people who order the tests.
  • It interprets assertively. Its readings of motive and pattern are stated as conclusions, not guesses, and they are usually right. Its first read can lean your way; ask for the other side and it will give it. If a reading doesn't fit, say so. It takes correction well.
  • Across separate sittings, tell it the date. It logs "it's been nine days" faithfully but keeps counting from the last date it knew. One line ("it's the 18th") fixes the arithmetic.
  • Long sessions may require occasional corrections.
  • Open weights, Apache 2.0. Deploy responsibly.

The Therapy Line

Model Size For Status
Therapy 3.8 (this model) 27B full-depth work, serious hardware this release
Therapy-27B 27B previous-generation 27B available
Therapy-9B 9B the everyday driver (~6–10 GB) available
Fable-Therapy-9B · 4B 9B/4B earlier generation available

Choosing Your Model

Model Best For
Therapy 3.8 (this model) The deepest sessions: interpretive work, record integrity under pressure, long arcs
Therapy-9B Same design on everyday hardware, strongest at focused sessions
Opus-Therapy-9B Sibling lineage, Opus-distilled disposition

Dataset

Not released.


Built by Verdugie, independent ML researcher · OpusReasoning@proton.me

Downloads last month
451
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Verdugie/Therapy-3.8

Base model

Qwen/Qwen3.8-27B
Quantized
(973)
this model