Por favor, use este identificador para citar o enlazar este ítem:
https://hdl.handle.net/20.500.12008/56413
Cómo citar
Registro completo de metadatos
| Campo DC | Valor | Lengua/Idioma |
|---|---|---|
| dc.contributor.author | Sastre, Ignacio | - |
| dc.contributor.author | Etcheverry, Lorena | - |
| dc.contributor.author | Rey, Guillermo | - |
| dc.contributor.author | Moncecchi, Guillermo | - |
| dc.contributor.author | Rosá, Aiala | - |
| dc.date.accessioned | 2026-08-19T17:52:48Z | - |
| dc.date.available | 2026-08-19T17:52:48Z | - |
| dc.date.issued | 2025 | - |
| dc.identifier.citation | Sastre, I., Etcheverry, L., Rey, G. y otros. Post-OCR correction using large language models with constrained decoding [Preprint] Publicado en: Researche Square, julio 2025. 16 p. DOI: doi.org/10.21203/rs.3.rs-6823036/v1. | es |
| dc.identifier.uri | https://hdl.handle.net/20.500.12008/56413 | - |
| dc.description.abstract | This article addresses the problem of correcting noisy Optical Character Recognition (OCR) outputs from digitized historical documents, specifically those from the Berrutti Archive related to Uruguay’s civic-military dictatorship. These documents—produced with typewriters, diverse layouts, and overlaid annotations—pose significant hallenges for standard OCR tools, resulting in highly errorprone text. We present a novel post-OCR correction method that leverages fine-tuned open-source Large Language Models (LLMs) combined with a constrained decoding strategy. This strategy incorporates character-level similarity between the OCR input and the generated output at decoding time, steering the model toward corrections that closely preserve the original text structure. We evaluate our method on a gold-standard dataset of over 2000 annotated lines and show that it outperforms prompting and standard fine-tuning approaches, reducing both character error rate (CER) and word error rate (WER). The corrected outputs provide more accurate input for downstream tasks, such as named entity recognition, relation and event extraction, and knowledge graph construction, thereby supporting the broader goal of extracting knowledge from historically significant and sensitive archives. | es |
| dc.description.sponsorship | Proyecto ANII IA_1_2022_1_173863 | es |
| dc.format.extent | 16 p. | es |
| dc.format.mimetype | application/pdf | es |
| dc.language.iso | en | es |
| dc.rights | Las obras depositadas en el Repositorio se rigen por la Ordenanza de los Derechos de la Propiedad Intelectual de la Universidad de la República.(Res. Nº 91 de C.D.C. de 8/III/1994 – D.O. 7/IV/1994) y por la Ordenanza del Repositorio Abierto de la Universidad de la República (Res. Nº 16 de C.D.C. de 07/10/2014) | es |
| dc.subject | Natural Language Processing | es |
| dc.subject | Optical Character Recognition (OCR) | es |
| dc.subject | Post-OCR Correction | es |
| dc.subject | LLMs with Constrained Decoding | es |
| dc.title | Post-OCR correction using large language models with constrained decoding | es |
| dc.type | Preprint | es |
| dc.contributor.filiacion | Sastre Ignacio, Universidad de la República (Uruguay). Facultad de Ingeniería. Instituto de Computación. | - |
| dc.contributor.filiacion | Etcheverry Lorena, Universidad de la República (Uruguay). Facultad de Ingeniería. Instituto de Computación. | - |
| dc.contributor.filiacion | Rey Guillermo, Universidad de la República (Uruguay). Facultad de Ingeniería. Instituto de Computación. | - |
| dc.contributor.filiacion | Moncecchi Guillermo, Universidad de la República (Uruguay). Facultad de Ingeniería. Instituto de Computación. | - |
| dc.contributor.filiacion | Rosá Aiala, Universidad de la República (Uruguay). Facultad de Ingeniería. Instituto de Computación. | - |
| dc.rights.licence | Licencia Creative Commons Atribución - No Comercial - Sin Derivadas (CC - By-NC-ND 4.0) | es |
| Aparece en las colecciones: | Publicaciones académicas y científicas - Instituto de Computación | |
Ficheros en este ítem:
| Fichero | Descripción | Tamaño | Formato | ||
|---|---|---|---|---|---|
| SERMR25.pdf | Preprint | 1,88 MB | Adobe PDF | Visualizar/Abrir |
Este ítem está sujeto a una licencia Creative Commons Licencia Creative Commons