Repository navigation
Conversation
…ed event log line reader Co-authored-by: Claude Code
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
The bounded line reader used by
ReplayListenerBus.replay(InputStream, ...)(added in SPARK-59407, byte accounting fixed in SPARK-59804) calledBufferedReader.read()once per character. This PR makes it read decoded characters in bulk instead:boundedLinesreads from theInputStreamReaderinto an 8192-char array and scans that array for\n, instead of wrapping the reader in aBufferedReaderand callingread()per character.BufferedReader.read()takes a lock on every call.BoundedLineBuffergetsappend(chars, start, end), which has the same result as appending the characters one by one: it counts UTF-8 bytes per character, stops counting at the first character that exceeds the allowance, and appends the range in one call if it fits or releases the retained prefix otherwise.append(c)is kept and shares the width and release helpers.The reader's semantics are unchanged: the UTF-8 content byte limit with the trailing-CR allowance, physical line indices including skipped lines, strict decoding, and releasing the buffer of an over-long line while it is drained. The
Iterator[String]overload ofreplayis not affected.Why are the changes needed?
Since SPARK-59407, line splitting in this reader is several times slower than the previous
Source.getLines()/readLine()(measured at about 14x in the review of #59074). Every replay through theInputStreamoverload pays this: loading application UIs in the History Server, event log compaction, and listing with fast in-progress parsing.It also matters for #59327 (SPARK-60109), which routes the History Server end-event reparse through this reader. For compressed logs, that reparse reads most of the decompressed file, because the skip target is computed from the compressed length, so without this change the listing of completed applications would become several times slower.
A local measurement of a full replay through
replay(InputStream, ...)with the History Server listing filter (210 MB of uncompressed event lines built fromspark-events/local-1642039451826with randomized digits, median of 3 runs, JDK 21):Source.getLines()+replay(Iterator)Does this PR introduce any user-facing change?
No behavior change. Replaying event logs through the bounded reader (History Server UI loading, listing and event log compaction) is faster.
How was this patch tested?
New tests in
ReplayListenerSuite:SPARK-60158: Line buffer range appends match single-character appends: randomized characters (ASCII, CR, 2- and 3-byte characters, surrogate halves), limits and range splits; the result, the retained length and the release of the buffer match appending one character at a time.SPARK-60158: Replay splits lines across read buffers like a reference splitter: randomized inputs with lines up to 20000 characters, so that lines and multi-byte characters cross the 8192-char read buffer, compared with a reference that splits on\n, drops one trailing CR and keeps lines within the byte limit.The existing SPARK-59804 tests (byte limit grid across the read buffer boundary, physical line numbers, skipped line diagnostics, malformed UTF-8 and stream IO errors) pass unchanged.
ReplayListenerSuite,FsHistoryProviderSuite(RocksDB backend),EventLogFileCompactorSuite,core/scalastyleandcore/Test/scalastylepass.Was this patch authored or co-authored using generative AI tooling?
Generated-by: Claude Code (Claude Opus 5.5)