VISHC at PsyDefDetect: Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation
Accepted at BioNLP @ ACL 2026 — PsyDefDetect Shared Task
Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Automatically classifying PDMs from text is clinically valuable but severely hindered by data scarcity and class imbalance — challenges that generative augmentation alone cannot resolve without psychological grounding.
We address these challenges in the PsyDefDetect shared task (BioNLP@ACL 2026) by proposing a context-aware synthetic augmentation framework combined with a hybrid classification model. Our hybrid model integrates contextual language representations with clinical features, alongside 150 annotated defense items. Experiments demonstrate that definition quality in prompting directly governs generation fidelity and downstream performance.
Our method surpasses the DMRS Co-Pilot baseline, establishing a strong benchmark for psychologically grounded defense mechanism classification in low-resource settings.
| Model | Accuracy | Macro-F1 |
|---|---|---|
| DMRS Co-Pilot (baseline) | 18.01% | 8.63% |
| CASA-PDC (ours) | 58.26% | 24.62% |
| Δ improvement | +40.25% | +15.99% |
Requires Python 3.10 and CUDA 12.x for GPU support. Adjust the CUDA version in requirements.txt if needed.
conda create -n psydef_env python=3.10 -y
conda activate psydef_env
pip install -r requirements.txtThis project uses the PsyDefDetect dataset from the BioNLP@ACL 2026 shared task.
-
Register and download the dataset from the official shared task page: 👉 PsyDefDetect Shared Task
-
Place the downloaded files in the
data/directory:
data/
├── train.json
└── test.json- Run the preprocessing script to prepare the data for training:
python scripts/preprocess.py --data_dir data/ --output_dir data/processed/Note: The dataset is provided solely for research purposes under the terms of the shared task organizers. Please review their data usage agreement before downloading.
This pipeline uses MentalRoBERTa as the backbone model. Access requires approval from the model owner on Hugging Face.
Setup your Hugging Face token:
huggingface-cli login
# or set the environment variable:
export HUGGINGFACE_TOKEN=your_token_hereRun training:
chmod +x script/run.sh
./script/run.shYou will be prompted to enter your Hugging Face API key if not already set. After successful authentication, training will begin on the PsyDefDetect dataset. Outputs and checkpoints are saved to the directory specified in script/run.sh.
For hyperparameter details and evaluation metrics, refer to script/run.sh and the corresponding sections in the paper.
For questions or issues, please contact: Hoang-Thuy-Duong Vu — 26duong.vht@vinuni.edu.vn
If you find this work useful, please cite our paper:
@inproceedings{vu-etal-2026-vishc,
title = "{VISHC} at {P}sy{D}ef{D}etect: Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation",
author = "Vu, Hoang-Thuy-Duong and
Pham, Quoc-Cuong and
Pham, Huy-Hieu",
editor = "Gupta, Deepak and
Demner-Fushman, Dina",
booktitle = "Proceedings of the {B}io{NLP} 2026 (Shared Tasks)",
month = jul,
year = "2026",
address = "San Diego, California, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.bionlp-2.12/",
doi = "10.18653/v1/2026.bionlp-2.12",
pages = "77--86",
ISBN = "979-8-89176-435-4",
abstract = "Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Automatically classifying PDMs from text is clinically valuable but severely hindered by data scarcity and class imbalance, challenges which generative augmentation alone cannot resolve without psychological grounding. In this work, we address these challenges in the PsyDefDetect shared task (BioNLP@ACL 2026) by proposing a context-aware synthetic augmentation framework combined with a hybrid classification model. Our hybrid model integrates contextual language representations with basic clinical features, along with 150 annotated defense items. Experiments demonstrate that definition quality in prompting directly governs generation fidelity and downstream performance. Our method surpasses DMRS Co-Pilot, reaching an accuracy of 58.26{\%} (+40.25{\%}) and a macro-F1 of 24.62{\%} (+15.99{\%}), thereby establishing a strong baseline for psychologically grounded defense mechanism classification in low-resource settings. Source code is available at: \url{https://github.com/htdgv/CASA-PDC}."
}