Refactor big-data processing trait, add chunkById(), remove Fiber variants - #343
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request Checklist
Summary
Fixes data-loss and performance problems in
InteractsWithBigDataProcessing(chunk,cursor,stream,batch), addschunkById(), and removesfchunk(),fstream()andfcursor().Bugs fixed
fstream()dropped the first item of every chunk and yieldednullafter each chunk.fcursor()never passed the first$bufferSizerows to the processor and raised aforeach(null)warning at the end.chunk(),stream()andbatch()usedOFFSETwith noORDER BY, so pages could overlap or skip rows. They now order by the primary key when the query has no ordering and noGROUP BY.limit()oroffset()on the query was overwritten, so->limit(100)->chunk(10, ...)scanned the whole table. Both are now respected.$processedinchunk()started at$chunkSize, even when the first page was smaller. It is now the running count including the current chunk.InvalidArgumentException.cursor()rethrew database errors without the originalPDOException. It now keeps it as the previous exception.Performance
cursor()no longer callsgc_collect_cycles()on every row.cursor()now hydrates rows throughfetchLazy(), so it matchesget()(connection name, encrypted attributes).New
chunkById($chunkSize, $processor, $column = null, $total = null)pages withWHERE id > :last ORDER BY id. It stays fast on very large tables, and it is safe when the processor updates or deletes rows that match theWHEREclause.OFFSETpaging skips rows in that case.cursor(..., bool $unbuffered = false)streams rows from the server on MySQL. By default the MySQL PDO driver loads the whole result set into client memory. It is opt-in because an unbuffered connection cannot run other queries while the cursor is open.Removed (breaking)
fchunk(),fstream()andfcursor()are removed. Fibers are cooperative and everything ran on one thread and one PDO connection, so these methods never ran in parallel, and the buffering gave no backpressure. Migration:fchunk($size, $processor, $concurrency)chunk($size, $processor)fstream($size, $transform, $bufferSize)stream($size, $transform)fcursor($processor, $bufferSize)cursor($processor)Behavior changes to note
chunk()andstream()may now addORDER BY <primary key>to the generated SQL.chunk(),chunkById(),stream()andbatch()throwInvalidArgumentExceptionfor a chunk size below 1.cursor()processor index is the 1-based row number.Docs
entity-orm.mdin the docs repo is updated:chunkById(), buffered vs. unbufferedcursor(), the removal table above, and the behavior changes. That change is in a separate repo and needs its own PR.Checklist