# Parallel gVCF import, possible?

**URL:** <https://forum.tiledb.com/t/parallel-gvcf-import-possible/816>\
**Category:** Uncategorized\
**Created:** [November 11, 2025, 1:10am UTC](https://forum.tiledb.com/t/parallel-gvcf-import-possible/816 "2025-11-11T01:10:21Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Michaeljon\_Miller](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.tiledb.com/michaeljon_miller/32/386_2.png) [@Michaeljon\_Miller](https://forum.tiledb.com/u/Michaeljon_Miller)\
**Post date:** [November 11, 2025, 1:10am UTC](https://forum.tiledb.com/t/parallel-gvcf-import-possible/816/1 "2025-11-11T01:10:21Z")

</div>

We run several hundred sequences through our variant calling pipeline every day. When a sequence passes automated QC a workflow is fired off to import the resulting gVCF into our s3 tiledb store. Is there any reason we can’t have many independent `tiledbvcf import` running concurrently against the same bucket? Or will that cause issues with the underlying data?

---

<div class="post-metadata">

**Author:** ![ihnorton](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.tiledb.com/ihnorton/32/23_2.png) [@ihnorton](https://forum.tiledb.com/u/ihnorton)\
**Post date:** [November 14, 2025, 1:13pm UTC](https://forum.tiledb.com/t/parallel-gvcf-import-possible/816/2 "2025-11-14T13:13:00Z")

</div>

Hi @Michaeljon_Miller,

As long as the sample IDs are distinct for each worker, then writing from multiple workers simultaneously is fine.

Best,

Isaiah

---

<div class="post-metadata">

**Author:** ![Michaeljon\_Miller](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.tiledb.com/michaeljon_miller/32/386_2.png) [@Michaeljon\_Miller](https://forum.tiledb.com/u/Michaeljon_Miller)\
**Post date:** [November 16, 2025, 10:12pm UTC](https://forum.tiledb.com/t/parallel-gvcf-import-possible/816/3 "2025-11-16T22:12:32Z")

</div>

Thanks Isaiah. This does seem to work, but I suspect that we’ll need to be a lot smarter about batching and consolidating. Running a few batches in serially (50 samples per import) seems to leave things in a good state. But I ran 40 concurrent batches of 50 and now read performance has gone south.

I tried to consolidate but ran into this.

```auto
[2025-11-16 20:04:45.612] [tiledb-vcf] [Process: 622] [Thread: 622] [critical]
 Exception: SparseIndexReaderBase: Cannot set array memory budget (3276.600000) 
because it is smaller than the current memory usage (19185324).

```

And that was with 64Gb assigned and 32 cores.

```bash
tiledbvcf utils consolidate fragments \
    --uri ${variant_store} \
    --tiledb-config sm.mem.total_budget=65536,sm.compute_concurrency_level=32 \
    --log-level trace \
    --log-file ${meta.nfr_workflow_id}.tiledb.log

```

Something tells me our use case might be straining the model a bit. Of course we want column-wise statistics (how many samples with variant x), but two other cases are to retrieve a single sample gVCF and to merge N samples into a single gVCF (where N is going to range up to, and beyond, 25,000 samples, but starting at 5,000 by the end of this year).

---

<div class="post-metadata">

**Author:** ![michaeljon](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.tiledb.com/michaeljon/32/274_2.png) [@michaeljon](https://forum.tiledb.com/u/michaeljon)\
**Post date:** [December 4, 2025, 2:30am UTC](https://forum.tiledb.com/t/parallel-gvcf-import-possible/816/4 "2025-12-04T02:30:08Z")

</div>

Still running into a OOM trying to run consolidation. How do I tell the CLI that it’s free to use the 192gb on the machine?

---

<div class="post-metadata">

**Author:** ![michaeljon](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.tiledb.com/michaeljon/32/274_2.png) [@michaeljon](https://forum.tiledb.com/u/michaeljon)\
**Post date:** [December 5, 2025, 5:10pm UTC](https://forum.tiledb.com/t/parallel-gvcf-import-possible/816/5 "2025-12-05T17:10:48Z")

</div>

Didn’t realize `sm.mem.total_budget` was in bytes. Either way, I’ve adjusted that but still running into OOM issues. I have a store with 250 samples, loaded in batches of 5x10 (50 into `tiledbvcf store` with a batch size of 10). Between each of those is a consolidation / vacuum pass on fragments and commits. I’ve also completely given up on using gVCF as the source and have moved to using our VCFs instead.

How much RAM should that consolidation take? And, generally, how much time if the store is on s3? I’m having a hard time seeing how this is going to scale into the 1000s let alone our target of 100000+.
