Applications
Domain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations
arXiv:2604.14815v1 Announce Type: new Abstract: In NLP classification tasks where little labeled data exists, domain fine-tuning of transformer models on unlabeled data is an established approach. In
arXiv:2604.14815v1 Announce Type: new Abstract: In NLP classification tasks where little labeled data exists, domain fine-tuning of transformer models on unlabeled data is an established approach. In this paper we have two aims. (1) We describe our observations from fine-tuning the Finnish BERT model on Finnish medical text data. (2) We report on our attempts to predict the benefit of domain-specific pre-training of Finnish BERT from observing the geometry of embedding changes due to domain fine-tuning. Our driving motivation is the commonsituation in healthcare AI where we might experience long delays in acquiring datasets, especially with respect to labels.
Related
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions
- Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
- Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection
- LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs
Source: arXiv cs.CL | 2026-04-17