Skip to content

Releases: aws-samples/amazon-textract-textractor

Version 1.10.0

Choose a tag to compare

@Belval Belval released this 11 Aug 13:58
8ea5f9a

What's Changed

  • Replace editdistance with rapidfuzz for Levenshtein distance by @Belval in #452
  • Fix undefined name 't2' in get_layout_csv_from_trp2 by @Belval in #453
  • Add Table.to_list() for dependency-free row export by @Belval in #454
  • Add optional boto3 config parameter to Textractor by @Belval in #455
  • Fix get_cells_by_type returning only COLUMN_HEADER cells by @Belval in #457
  • Declare and test support for Python 3.13 and 3.14 by @Belval in #456
  • Fix failing Documentation CI deploy by @Belval in #458
  • Version 1.10.0 by @Belval in #459

Full Changelog: v1.9.2...v1.10.0

Version 1.9.2

Choose a tag to compare

@Belval Belval released this 24 Apr 14:14

What's Changed

Full Changelog: v1.9.1...v1.9.2

Version 1.9.1

Choose a tag to compare

@Belval Belval released this 27 Mar 20:59

What's Changed

  • Fix s3 client instantiation in _get_document_images_from_path by @Belval in #423

Full Changelog: v1.9.0...v1.9.1

Version 1.9.0

Choose a tag to compare

@Belval Belval released this 07 Mar 21:40

What's Changed

Full Changelog: v1.8.5...v1.9.0

Version 1.8.5

Choose a tag to compare

@Belval Belval released this 13 Nov 14:55

What's Changed

  • Fix bug in convert that caused an exception on empty pages.

Full Changelog: v1.8.4...v1.8.5

Version 1.8.4

Choose a tag to compare

@Belval Belval released this 06 Nov 22:29

What's Changed

  • Add check for None bounding boxes for AnalyzeExpense by @Belval
  • Allow Custom Separator in Document.export_kv_to_csv() by @Chuukwudi
  • Update analyze_document type hint by @ryangamble
  • Fix invalid escape in BoundingBox docstring by @simonschmidt in #395

Full Changelog: v1.8.3...v1.8.4

Version 1.8.3

Choose a tag to compare

@Belval Belval released this 21 Aug 16:22

What's Changed

  • Id in html output by @Belval in #386
  • Escape html output by @Belval in #387
  • Fix table indexing returning too many cells

⚠️ Breaking changes

  • To support ids in HTML, layout Table created for TABLE predictions will no longer share the same ID as the table.

Full Changelog: v1.8.2...v1.8.3

Version 1.8.2

Choose a tag to compare

@Belval Belval released this 25 Jun 13:44

What's Changed

  • Fix pypdfium2 failing to parse PDFs in bytearray format by @Belval

Full Changelog: v1.8.1...v1.8.2

Version 1.8.1

Choose a tag to compare

@Belval Belval released this 24 Jun 13:25

What's Changed

  • Fix .to_markdown() raising an exception on missing local config by @Belval in #381

Full Changelog: v1.8.0...v1.8.1

Version 1.8.0

Choose a tag to compare

@Belval Belval released this 21 Jun 00:54

What's Changed

  • Improve HTML linearization
    • Add HTML table linearization format that uses merged cells information for colspan and rowspan
    • Add prefix and suffix for LAYOUT_FOOTER and LAYOUT_ENTITY
    • Add <html><body>...</body></html> to the output when calling Document.to_html()
  • Use pypdfium2 for PDF rasterization when available instead of pdf2image. This allows for better portability as the former does not have a dependency on OS libraries and should work out of the box with Lambda and SageMaker.
  • Fix expenses with no summary fields
  • Replace region mismatch with invalid S3 object exception

Backward-incompatible changes

  • This update removes s3_output_path from the synchronous functions as s3_output_path is not a supported parameter for the Textract Synchronous API
  • This update changes the exception raised by the textractor.py functions which will no longer raise RegionMismatchError (which is however kept in textractor.exceptions for backward compatibility.
  • This update removes confidence_score from KeyValue entities in favour of _confidence which is used for all other entities.

Full Changelog: v1.7.12...v1.8.0