Stream: troubleshooting

Topic: File upload questions


view this post on Zulip Bethany Seeger (Aug 24 2026 at 15:40):

Hello,

We had a large upload over the weekend and our server misbehaved because of it. We've since reduced our upload limits to help with these situations til we figure this out, but have a few questions.

1) does DV it try to upload all files at once, or queue them in the case that there are 1000 files in one upload? We are wondering how many connections the server needs to have available to service uploads. Many of these files were 3-4 GB tif files -- perhaps thumbnail processing caused a bottleneck?

2) this dataset (still draft) has 154,000 files and the dataset page loads slowly (30s). Is the slowness mostly likely due to Solr? We are not seeing a huge CPU load when loading the page. Just wondering what is going on behind the scenes and how to optimize this.

We probably need to fine tune things, and just need to figure out what. Thanks for any help!

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 24 2026 at 15:47):

154,000 files?! Um. That's a lot! I owe you a better reply but for now, please take a look at https://guides.dataverse.org/en/6.11/admin/big-data-administration.html#avoiding-many-files

I'm wondering if you can store the files as a zip and use the zip previewer/downloader. Also from that page:

Dataverse can be configured to use a โ€œZip File Previewerโ€ that allows users to see the contents of a zip file and even download individual files from within it (seeย Compressed Files).

view this post on Zulip Bethany Seeger (Aug 24 2026 at 15:49):

Thanks, Philip. I'm not sure if they uploaded them in batches - hopefully they did.

view this post on Zulip Bethany Seeger (Aug 24 2026 at 19:30):

I am curious -- how often, in your experience, have you run across a super large number of files in a dataset?

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 24 2026 at 19:33):

Hmm, it's a better question for @Leo Andreev. He tells us about them. Pretty often, I think! But this setting should help:

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 24 2026 at 19:35):

It looks like Harvard Dataverse allows 1000 files:

Screenshot 2026-08-24 at 3.34.47โ€ฏPM.png

view this post on Zulip Bethany Seeger (Aug 24 2026 at 19:37):

Thanks! I changed ours this morning to be similar, so that will help some.

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 24 2026 at 19:38):

It looks like it's possible to set a total number of files via API: https://guides.dataverse.org/en/6.11/api/native-api.html#storage-quotas-on-individual-datasets

I'm seeing "numberOfFilesRemaining": 20,

view this post on Zulip Bethany Seeger (Aug 26 2026 at 20:04):

Hi Philip - is the Harvard system also limited to 1000 files in a zip file? So it's 1000 across the board, either zipped up or through drag/drop?

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 26 2026 at 20:09):

Well, the settings says this...

... which means it isn't relevant for Harvard Dataverse, which uses S3 direct upload.

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 26 2026 at 20:11):

So in practice, I'm pretty sure a zip can have an unlimited number of files.

view this post on Zulip Philip Durbin ๐Ÿš€ (Aug 27 2026 at 20:44):

Bethany Seeger said:

1) does DV it try to upload all files at once, or queue them in the case that there are 1000 files in one upload?

@Bethany Seeger are these uploads happening via the web interface? Or via API?

view this post on Zulip Bethany Seeger (Sep 02 2026 at 20:54):

They are happening via the web interface, I believe - via a zip file.

view this post on Zulip Philip Durbin ๐Ÿš€ (Sep 03 2026 at 11:22):

Interesting. So the zips files are unzipped, then, resulting in lots and lots of files?


Last updated: Oct 02 2026 at 18:51 UTC