An Azure service that turns documents into usable data. Previously known as Azure Form Recognizer.
Hi Marco Nijmeijer,
We discussed your latest findings with our Backend team. Based on their further investigation, they have confirmed that this is a character recognition issue with the Document Intelligence OCR/Read model, specifically affecting the recognition of the “ë” character in certain scenarios.
The PG team clarified that, for this particular recognition issue, enabling High Resolution or Language Detection does not address the underlying behavior. Therefore, the inconsistent recognition of “ë” at different font sizes can still occur even when Dutch (nl) is specified.
At this time, the PG team recommends trying Azure Content Understanding, where improvements have been made to the recognition algorithms. This may provide better recognition of characters with diacritics such as “ë” and could help address the behavior you are experiencing.
Could you please test the same document using Content Understanding and let us know whether the character is recognized correctly across the different font sizes?
We are also continuing to validate the behavior against the latest available version to ensure there are no regressions.
We appreciate your detailed testing and feedback. Please share the results of the Content Understanding test, and we will continue to work with the PG team based on the outcome.
We tested the latest AR model, and the results look good. No regression issues were observed during testing.