davanstrien's picture
davanstrien HF Staff
Add model card with training details and metrics
961aaba verified
|
Raw
History Blame Contribute Delete
2.97 kB
metadata
license: apache-2.0
base_model: answerdotai/ModernBERT-base
datasets:
  - biglam/on_the_books
language:
  - en
pipeline_tag: text-classification
tags:
  - legal
  - classification
  - jim-crow
  - modernbert
metrics:
  - f1
  - accuracy
  - precision
  - recall
model-index:
  - name: jim-crow-laws-ml-agent
    results:
      - task:
          type: text-classification
          name: Text Classification
        dataset:
          name: biglam/on_the_books
          type: biglam/on_the_books
          split: test
        metrics:
          - name: F1
            type: f1
            value: 0.9487
          - name: Accuracy
            type: accuracy
            value: 0.9701
          - name: Precision
            type: precision
            value: 0.9367
          - name: Recall
            type: recall
            value: 0.961

Jim Crow Law Classifier (ModernBERT-base)

A text classification model fine-tuned on biglam/on_the_books to identify Jim Crow laws in historical US legislative text.

Model Description

This model classifies sections of US state legislation as either Jim Crow laws (discriminatory laws targeting racial minorities) or non-Jim Crow laws. It was fine-tuned from answerdotai/ModernBERT-base, which supports up to 8,192 tokens of context.

Performance

Evaluated on a stratified 15% held-out test set (268 samples):

Metric Score
F1 0.9487
Accuracy 0.9701
Precision 0.9367
Recall 0.9610

Training Details

  • Base model: answerdotai/ModernBERT-base (149M parameters)
  • Dataset: biglam/on_the_books (1,785 samples total; 1,517 train / 268 test)
  • Max sequence length: 1024 tokens
  • Epochs: 5 (best checkpoint at epoch 5 by F1)
  • Batch size: 16
  • Learning rate: 2e-5 with linear decay
  • Warmup: 6% of training steps
  • Weight decay: 0.01
  • Hardware: NVIDIA T4 GPU
  • Training time: ~8 minutes

Usage

from transformers import pipeline

classifier = pipeline("text-classification", model="davanstrien/jim-crow-laws-ml-agent")

text = "The Commission shall provide separate sleeping quarters and separate eating space for the different races."
result = classifier(text)
print(result)
# [{'label': 'jim_crow', 'score': 0.99...}]

Dataset

The On the Books dataset contains 1,785 sections of North Carolina state legislation from the Jim Crow era, annotated by historians as either Jim Crow laws or non-Jim Crow laws. The dataset is imbalanced: 71% non-Jim Crow, 29% Jim Crow.

Labels

  • no_jim_crow (0): Non-discriminatory legislation
  • jim_crow (1): Jim Crow law (racially discriminatory legislation)

Limitations

  • Trained only on North Carolina legislation; may not generalize to other states
  • Historical language patterns may not transfer to modern legal text
  • The model may be biased toward the specific annotation criteria used in the dataset