Skip to content

When pushing a blob stream, compute digest on demand #278

Description

@lehmanju

The current implementation of push_blob_stream requires specifying the digest up front. For streams that process bytes on demand, this digest can only be fully computed once the stream has been consumed. Therefore, I'd like to use something like this:

pub async fn push_blob_stream<T: Stream<Item = Result<bytes::Bytes>> + Send + 'static>(
        &self,
        image: &Reference,
        blob_data_stream: T,
        // blob_digest: &str, -- REMOVE
    ) -> Result<String> {
        let mut location = self.begin_push_chunked_session(image).await?;
        let mut range_start = 0;

        let mut blob_data_stream = pin!(blob_data_stream);

        while let Some(blob_data) = blob_data_stream.next().await {
            // compute hash
            let mut blob_data = blob_data?;
            while !blob_data.is_empty() {
                let chunk = blob_data.split_to(self.push_chunk_size.min(blob_data.len()));
                (location, range_start) = self
                    .push_chunk(&location, image, chunk, range_start)
                    .await?;
            }
        }
        let blob_digest = format!(..., hash) // NEW, compute digest on demand
        self.end_push_chunked_session(&location, image, blob_digest)
            .await
    }

My use case:

Download large files from a remote HTTP location, put each file into a separate oci layer and push a container image to registry. Since the files are large and are downloaded anyway, this would ideally skip local caching and connect the streams directly.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions