Sabine Richling, Sven Siebler, Alexander Balz, Robert Kühl, Martin Baumann
5 min
SDS@hd is a centralized scientific data storage service operated by the Heidelberg University Computing Centre. Launched in 2017, it provides a high-performance, secure storage backend tailored for the active phases of the research data life cycle. It is designed to meet the needs of researchers across all universities in Baden-Württemberg who generate large-scale data through high-throughput instruments, simulations, or compute-intensive workflows.
The service functions as a collaborative workspace where researchers can manage data through defined storage projects. Access is managed via the federated identity management system (bwIDM), allowing users to authenticate using their home institution credentials. The system supports fine-grained access control, including role-based permissions and filesystem-level access control lists (ACLs), which facilitate secure data sharing among project members and external collaborators.
SDS@hd is specifically engineered for performance, offering direct, high-speed connectivity to the local high-performance compute cluster (bwForCluster MLS&WISO). This integration is critical for researchers in fields like life sciences, medical imaging, and astrophysics, who require rapid access to large datasets during simulation, post-processing, and analysis.
While SDS@hd excels at supporting active research, it is not a repository for long-term preservation or open-access publication. The authors emphasize that SDS@hd is part of a broader ecosystem; researchers are expected to transfer finalized data to dedicated platforms—such as heiDATA for publication or heiARCHIVE for long-term storage—to ensure metadata standards and persistent identification are maintained.
As research data volumes continue to grow due to advancements in high-throughput technology, institutional storage solutions that bridge the gap between raw data generation and final publication are essential. By providing a scalable, collaborative, and compute-integrated environment, SDS@hd reduces the technical burden on individual research groups and supports the broader goal of data-intensive science within the state of Baden-Württemberg.
Alex: [deliberate, checking understanding] So the system abstracts away both the throughput problem and the identity problem. What happens once a project wraps up and the data needs to move into long-term preservation? [[RP_SECTION:long-term-preservation|Long Term Preservation]]
Sam: [direct, acknowledging the limitation] That's the real constraint, and the authors are upfront about it. SDS@hd is strictly hot storage. It has no native metadata curation and no automated FAIR-compliance tooling built in. Researchers have to manually migrate their data out to external archives—heiDATA or heiARCHIVE—when a project ends. It's a high-performance bucket, not a self-describing data fabric.
Alex: [slower, reflective] Is there any indication they're planning to close that gap? [[RP_SECTION:future-infrastructure-evolution|Future Infrastructure Evolution]]
Sam: [thoughtful] The paper points to future work on federating data across the state's various heterogeneous services, and on building automated metadata extraction pipelines. If that lands, it would shift the system from passive storage toward something closer to an active research data fabric—but that's forward-looking, not something the current deployment does.
Alex: [pace picking up, connecting the dots] That would remove a real friction point. Right now the researcher is the one responsible for shepherding data from hot storage into a properly curated archive.
Sam: [nodding in voice, precise] Exactly. The system is genuinely effective at the compute-heavy phase of the lifecycle. The transition to long-term storage remains a manual step, and that's a deliberate scope decision rather than an oversight—they built for the most acute bottleneck first.
Alex: [deliberate, summarizing] So the value proposition is throughput and federated access, with lifecycle management still resting on the researcher's discipline.
Sam: [quiet confidence] That's a fair read. The evidence for it working is the adoption pattern—uptake across fields as different as the life sciences and digital humanities, which suggests the architecture generalizes beyond any single discipline's workflow. The open question going forward is whether they can scale the federation and metadata layer without eroding the performance that makes the hot tier worth using in the first place.
Alex: [reflective] It's a good illustration of how much the physical constraints of data generation end up shaping the architecture of shared research infrastructure.
Sam: [warm, professional, concluding] It's a measured, incremental design choice—solve the immediate pain of data gravity threatening to stall high-throughput instruments, and leave the harder, system-wide problem of long-term metadata curation for a later phase of the project.
Alex: [settling] If you want the specifics on the access control implementation and the migration workflows we didn't fully unpack, you can generate a deep dive of this paper. The report itself has the rest either way.
Sam: Thanks for listening.