Skip to content

Skip articles that haven't changed between dumps #9

Open
@newsch

Description

The dump schema includes a date_modified timestamp and other revision metadata.

To reduce disk I/O, we could store some metadata along the articles, compare it against the new one when processing, and skip them if they haven't changed.

One way to do this would be to store the date_modified timestamp as the modified attribute of the article file.

Activity

biodranik

biodranik commented on Jun 26, 2023

@biodranik
Member

An interesting optimization, but it may not worth it. Need to prove its benefits first. Let's leave it in a very low priority for now.

newsch

newsch commented on Jun 26, 2023

@newsch
CollaboratorAuthor

Understood, I've been thinking of it since you mentioned it here, we'll see what the profiling shows for the workflow.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

      Participants

      @biodranik@newsch

      Issue actions

        Skip articles that haven't changed between dumps · Issue #9 · organicmaps/wikiparser