Text Classification
Scikit-learn
Joblib
English
intent-classification
input-guard
medical-triage
pre-triage
scikit-learn
tf-idf
logistic-regression
healthcare
symptom-checker
academic-project
Eval Results (legacy)
Instructions to use cristian-untaru/sortmed-intent-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use cristian-untaru/sortmed-intent-classifier with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("cristian-untaru/sortmed-intent-classifier", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
File size: 10,639 Bytes
5ee09bf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 | ---
license: mit
language:
- en
tags:
- intent-classification
- input-guard
- medical-triage
- pre-triage
- text-classification
- scikit-learn
- tf-idf
- logistic-regression
- healthcare
- symptom-checker
- academic-project
pipeline_tag: text-classification
library_name: sklearn
model-index:
- name: sortmed-intent-classifier
results:
- task:
type: text-classification
name: Input intent classification for medical pre-triage
metrics:
- type: accuracy
value: 0.875
name: Accuracy
- type: f1
value: 0.8558
name: Macro F1
- type: precision
value: 0.8734
name: Macro Precision
- type: recall
value: 0.8634
name: Macro Recall
---
# SortMed Intent Classifier
## Model Description
This repository contains the intent classifier used by the **SortMed** input guard. The model decides whether a user input is a valid symptom description or whether it belongs to an unsafe or out-of-scope intent category before the text is sent to the medical triage models.
The classifier is a lightweight scikit-learn pipeline based on:
- TF-IDF text vectorization;
- Logistic Regression multi-class classification.
It is intentionally separate from the triage models. The triage models predict urgency only after this intent classifier and the deterministic input-guard rules accept the input.
This model was developed as part of the **SortMed** academic project, a medical pre-triage assistant prototype built for a bachelor's thesis by **Cristian Untaru** at the **West University of Timisoara, Faculty of Informatics**.
## Role in SortMed
The final SortMed input validation flow is:
```text
User input
|
v
Deterministic input-guard rules
|
v
Intent classifier
|
v
Triage model, only if intent = symptom_description
```
The classifier is used as a semantic safety layer. It blocks prompts that may contain medical words but are not suitable symptom descriptions, such as medication requests, diagnosis requests, general medical questions, or non-medical input.
## Intended Use
This model is intended to be used in the SortMed academic prototype for:
- classifying user input intent before medical pre-triage;
- rejecting unsafe or out-of-scope requests;
- allowing only English symptom descriptions to reach the triage classifier;
- supporting a hybrid input-guard architecture based on rules plus intent classification.
Example accepted input:
```text
I have chest pain and I feel short of breath.
```
Expected intent:
```text
symptom_description
```
Example rejected input:
```text
Can you recommend a painkiller for my headache?
```
Expected intent:
```text
medication_request
```
## Out-of-Scope Use
This model must not be used as:
- a medical triage classifier;
- a diagnostic model;
- a medication recommendation system;
- a replacement for deterministic safety rules;
- a standalone medical safety system;
- a general-purpose moderation classifier;
- a multilingual intent classifier without additional validation.
The model only classifies intent. It does not assess symptom severity and does not provide medical advice.
## Intent Classes
| Intent class | Meaning | SortMed behavior |
|---|---|---|
| `symptom_description` | The user describes symptoms or how they feel. | Accepted for triage if the confidence is high enough. |
| `medication_request` | The user asks for medication, drugs, treatment, or dosage advice. | Rejected with a medication-specific safety message. |
| `diagnosis_request` | The user asks directly for a diagnosis or condition identification. | Rejected with a diagnosis-specific safety message. |
| `general_medical_question` | The user asks a general medical question instead of describing symptoms. | Rejected with a message asking for a symptom description. |
| `non_medical` | The input is unrelated to medical symptoms. | Rejected as out of scope. |
Only `symptom_description` is considered a valid intent for continuing to the triage models.
## Configuration
The published configuration is stored in [`intent_config.json`](./intent_config.json).
| Field | Value |
|---|---:|
| Model type | `tfidf+logreg` |
| Number of classes | 5 |
| Valid triage intent | `symptom_description` |
| General confidence threshold | 0.5 |
| TF-IDF feature count | 1936 |
| Training examples | 382 |
| Test examples | 96 |
| scikit-learn version | `1.6.1` |
In the SortMed backend, `symptom_description` is accepted only when it passes the valid-intent confidence threshold used by the input guard. Other intents are rejected with class-specific user-facing messages.
## Evaluation Results
The classifier was evaluated on the held-out test split.
| Metric | Test Score |
|---|---:|
| Accuracy | 0.8750 |
| Macro Precision | 0.8734 |
| Macro Recall | 0.8634 |
| Macro F1 | 0.8558 |
The confusion matrix is available in [`intent_confusion_matrix.png`](./intent_confusion_matrix.png).
## How to Use
```python
from huggingface_hub import hf_hub_download
import joblib
import json
repo_id = "cristian-untaru/sortmed-intent-classifier"
model_path = hf_hub_download(repo_id=repo_id, filename="intent_pipeline.joblib")
config_path = hf_hub_download(repo_id=repo_id, filename="intent_config.json")
pipeline = joblib.load(model_path)
with open(config_path, "r", encoding="utf-8") as file:
config = json.load(file)
text = "I have chest pain and I feel short of breath."
probabilities = pipeline.predict_proba([text])[0]
classes = list(pipeline.classes_)
best_index = probabilities.argmax()
intent = classes[best_index]
confidence = float(probabilities[best_index])
print("Intent:", intent)
print("Confidence:", round(confidence, 4))
print("Valid intent:", config["valid_intent"])
```
Security note: `joblib` files rely on Python pickle serialization. Load this artifact only from trusted sources.
## Repository Files
| File | Description |
|---|---|
| `intent_pipeline.joblib` | Serialized scikit-learn TF-IDF + Logistic Regression pipeline. |
| `intent_config.json` | Intent classes, accepted intent, thresholds, feature count, split sizes, and scikit-learn version. |
| `intent_test_metrics.json` | Held-out test metrics for the intent classifier. |
| `intent_confusion_matrix.png` | Confusion matrix image for the intent classification task. |
| `README.md` | Model card documentation. |
| `.gitattributes` | Git LFS configuration for large files. |
## Limitations
This model has several important limitations:
- It is a TF-IDF + Logistic Regression classifier, not a contextual transformer or LLM.
- It may be sensitive to wording, spelling, unusual phrasing, or adversarial inputs.
- It was trained for English input only.
- It does not perform medical triage or diagnosis.
- It does not detect emergency severity.
- It should be used together with deterministic input validation rules.
- It should not be treated as a standalone safety system.
## Ethical and Safety Considerations
Input guards for medical applications must be conservative. This classifier is designed to reduce unsafe routing into the triage models, but it cannot guarantee perfect rejection of every invalid prompt.
For this reason, SortMed uses a hybrid validation design:
- deterministic rules for empty, repetitive, non-English, malformed, or adversarial input;
- this intent classifier for semantic request type detection;
- triage models only after the input is accepted as a symptom description.
Any production medical system would require clinical review, larger safety testing, monitoring, and a stronger risk-management process.
## Medical Disclaimer
This model is part of an academic prototype. It does not provide medical advice, diagnosis, treatment, or emergency triage.
If symptoms are severe, sudden, worsening, or potentially life-threatening, users should contact emergency services or a qualified healthcare professional immediately.
## Related SortMed Resources
### Triage Models
### Full Fine-Tuned Models
- [`cristian-untaru/distilbert-medical-triage`](https://huggingface.co/cristian-untaru/distilbert-medical-triage)
- [`cristian-untaru/biobert-medical-triage`](https://huggingface.co/cristian-untaru/biobert-medical-triage)
- [`cristian-untaru/roberta-medical-triage`](https://huggingface.co/cristian-untaru/roberta-medical-triage)
- [`cristian-untaru/biomedbert-medical-triage`](https://huggingface.co/cristian-untaru/biomedbert-medical-triage)
### LoRA Models
- [`cristian-untaru/lora-distilbert-medical-triage`](https://huggingface.co/cristian-untaru/lora-distilbert-medical-triage)
- [`cristian-untaru/lora-biobert-medical-triage`](https://huggingface.co/cristian-untaru/lora-biobert-medical-triage)
- [`cristian-untaru/lora-roberta-medical-triage`](https://huggingface.co/cristian-untaru/lora-roberta-medical-triage)
- [`cristian-untaru/lora-biomedbert-medical-triage`](https://huggingface.co/cristian-untaru/lora-biomedbert-medical-triage)
### Bottleneck MLP Adapter Models
- [`cristian-untaru/bottleneck-mlp-distilbert-medical-triage`](https://huggingface.co/cristian-untaru/bottleneck-mlp-distilbert-medical-triage)
- [`cristian-untaru/bottleneck-mlp-biobert-medical-triage`](https://huggingface.co/cristian-untaru/bottleneck-mlp-biobert-medical-triage)
- [`cristian-untaru/bottleneck-mlp-roberta-medical-triage`](https://huggingface.co/cristian-untaru/bottleneck-mlp-roberta-medical-triage)
- [`cristian-untaru/bottleneck-mlp-biomedbert-medical-triage`](https://huggingface.co/cristian-untaru/bottleneck-mlp-biomedbert-medical-triage)
### Frozen Encoder Models
- [`cristian-untaru/frozen-encoder-distilbert-medical-triage`](https://huggingface.co/cristian-untaru/frozen-encoder-distilbert-medical-triage)
- [`cristian-untaru/frozen-encoder-biobert-medical-triage`](https://huggingface.co/cristian-untaru/frozen-encoder-biobert-medical-triage)
- [`cristian-untaru/frozen-encoder-roberta-medical-triage`](https://huggingface.co/cristian-untaru/frozen-encoder-roberta-medical-triage)
- [`cristian-untaru/frozen-encoder-biomedbert-medical-triage`](https://huggingface.co/cristian-untaru/frozen-encoder-biomedbert-medical-triage)
### Datasets
- [`cristian-untaru/symcat-medical-triage-dataset`](https://huggingface.co/datasets/cristian-untaru/symcat-medical-triage-dataset)
- [`cristian-untaru/medquad-retrieval-pretriage`](https://huggingface.co/datasets/cristian-untaru/medquad-retrieval-pretriage)
|