Multi-label Classification and Named Entity Recognition for Historical Documents
| dc.contributor.author | Gruber, Ivan | |
| dc.contributor.author | Hlaváč, Miroslav | |
| dc.contributor.author | Neduchal, Petr | |
| dc.contributor.author | Hrúz, Marek | |
| dc.date.accessioned | 2026-03-30T18:05:35Z | |
| dc.date.available | 2026-03-30T18:05:35Z | |
| dc.date.issued | 2025 | |
| dc.date.updated | 2026-03-30T18:05:35Z | |
| dc.description.abstract | In this paper, we present improvements to our processing pipeline for historical document digitization. The original pipeline is extended with two new functionalities - page labeling, and named entity recognition. We handle page labeling as a multi-label classification task, for which we choose the Query2Label approach. Query2Label is tested on our internal NKVD dataset and reaches a mean average precision equal to 80.03% on the test set. For the named entity recognition task we utilize pre-trained transformer-based models DeepPavlov and benchmark them on two entities - person name, and location. The best model reaches promising results despite not being trained on our data at all. | en |
| dc.format | 11 | |
| dc.identifier.document-number | 001534826000002 | |
| dc.identifier.doi | 10.1007/978-3-031-81010-7_2 | |
| dc.identifier.isbn | 978-3-031-81009-1 | |
| dc.identifier.issn | 0302-9743 | |
| dc.identifier.obd | 43944197 | |
| dc.identifier.orcid | Gruber, Ivan 0000-0003-2333-433X | |
| dc.identifier.orcid | Hlaváč, Miroslav 0000-0003-1172-930X | |
| dc.identifier.orcid | Neduchal, Petr 0000-0001-5788-604X | |
| dc.identifier.orcid | Hrúz, Marek 0000-0002-7851-9879 | |
| dc.identifier.uri | http://hdl.handle.net/11025/67464 | |
| dc.language.iso | en | |
| dc.project.ID | DH23P03OVV073 | |
| dc.publisher | Springer | |
| dc.relation.ispartofseries | 7th International Conference on the Dynamics of Information Systems, DIS 2024 | |
| dc.subject | multi-label classification | en |
| dc.subject | named entity recognition | en |
| dc.subject | historical documents | en |
| dc.title | Multi-label Classification and Named Entity Recognition for Historical Documents | en |
| dc.type | Stať ve sborníku (D) | |
| dc.type | STAŤ VE SBORNÍKU | |
| dc.type.status | Published Version | |
| local.files.count | 1 | * |
| local.files.size | 964848 | * |
| local.has.files | yes | * |
| local.identifier.eid | 2-s2.0-105031157129 |
Files
Original bundle
1 - 1 out of 1 results
No Thumbnail Available
- Name:
- gruber_Multi-label Classification.pdf
- Size:
- 942.23 KB
- Format:
- Adobe Portable Document Format
License bundle
1 - 1 out of 1 results
No Thumbnail Available
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: