Skip to content

Ruby: Preserve Unicode character positions in Ruby extraction - #2783

Open
joelhawksley wants to merge 4 commits into
marcoroth:mainfrom
joelhawksley:joelhawksley-preserve-unicode-character-positions
Open

joelhawksley wants to merge 4 commits into
marcoroth:mainfrom
joelhawksley:joelhawksley-preserve-unicode-character-positions

Conversation

@joelhawksley

@joelhawksley joelhawksley commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

While working on #2260, I realized that some of the complexity in the implementation came from RuboCop operating on character offsets vs. Herb operating on byte offsets. I think for the public API we can mask non-Ruby text by Unicode character rather than UTF-8 byte.

This PR makes the following changes:

  • make preserve_positions: true mask non-Ruby text by Unicode character rather than UTF-8 byte
  • retain internal byte-preserving extraction for Prism analyzer offsets
  • document the contract and add cross-binding regression tests

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation ruby Ruby source for the gem and its libraries typescript TypeScript source across the javascript/ packages c C source for the core parser, lexer, and AST node @herb-tools/node native Node.js addon node-wasm @herb-tools/node-wasm WebAssembly parser for Node.js rubygem The herb RubyGem and its packaging rust Rust bindings and the Herb Rust crate java Java bindings and the org.herb package labels Oct 6, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@joelhawksley
joelhawksley marked this pull request as ready for review October 6, 2026 17:47
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added the rbs RBS type signatures in sig/ label Oct 6, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@marcoroth marcoroth changed the title Preserve Unicode character positions in Ruby extraction Ruby: Preserve Unicode character positions in Ruby extraction Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

c C source for the core parser, lexer, and AST documentation Improvements or additions to documentation java Java bindings and the org.herb package node @herb-tools/node native Node.js addon node-wasm @herb-tools/node-wasm WebAssembly parser for Node.js rbs RBS type signatures in sig/ ruby Ruby source for the gem and its libraries rubygem The herb RubyGem and its packaging rust Rust bindings and the Herb Rust crate typescript TypeScript source across the javascript/ packages

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant