The full vm on which gitlab.kitware.com is backed up every 12 hours to our local storage in the server room. These are backed up to a raided storage array with immutable storage. We currently have 2 backups per day going back 10 days. Plus we keep an offsite/offline copy of gitlab.kitware.com which gets updated monthly.
Is there a plan to mirror or archive the full issue, merge request, and discussion history (along with this discourse) somewhere broadly and reliably accessible?
Not currently. Any suggestions? Is this done for cloud hosted projects? Or does people rely on the backup strategy of the cloud host? Or is the main issue the “broadly and reliably accessible” part?
Yesterday I tried to open a normal issue from an email notification. After more than 30 minutes, the page still would not load. Instead, I was repeatedly presented with the Anubis proof-of-work challenge shown in the attached screenshots. If this were a one-off occurrence, I’d think nothing of it, but this has happened to me what feels like every time I’ve tried to access the GitLab for the last 6 months.
Working on getting answer to this from our sysadmin folks.
I got an answer from our sysadmin team. It is a long one so I will give a short summary here. Bottom line is that our infrastructure continues to evolve to fend off denial of service attacks. As an example, cmake.org was recently subjected to a massive DoS attack which ended up bringing down almost our entire server infrastructure. This was over the last few days and we changed how we handle Anubis as a result. So it is an evolving situation and let’s have patience to see if it settles.
As I mentioned before, we always have the option of moving to github while maintaining our current testing infrastructure (which was not available when we moved to gitlab).
Please keep reporting issues with the infrastructure here rather than just being frustrated and making voodoo dolls to curse the Kitware team. We have the same goals so let’s work together towards a solution.
@berk.geveci, thank you for forking the thread over here… that other thread went south quick and I’m regretful my proposal, attitude, or phrasing triggered that… My intention in all of this is to ensure there is a “backup” in place and that we as a community have a plan for any sort of disastrous scenario. My goal is not to see VTK move to GitHub (though I am confused why gitlab.com isn’t getting nearly as much attention in that other thread…)
Indeed this is my primary concern. It makes me incredibly nervous that all of this rich context, history, and community contributions are self-hosted and backed up all in ~one server room… Your note that it was almost entirely brought down by one DoS attack only adds to my nervousness. From the outside here it feels like the posterity of the project is at the whims of bad voodoo or a fault in the HVAC/plumbing systems.
My hot take here is that Kitware isn’t in the business of building data centers and cloud hosting… With millions of people depending on VTK, this doesn’t feel appropriate to self-host.
Someone brought up “the Iron Mountain incident” in the other thread as an example of how cloud providers are also vulnerable… My point here is that we’re all vulnerable. However, there are bigger initiatives and systems of redundancy built into these cloud providers that we can lean on and benefit from with the added benefit of reducing the strain Kitware’s resources.
I’ll explore some backup options and what other open source communities do to ensure posterity and reliable access. Would you please try to get a rough estimate of the size of these backups? I’m assuming its on the order of 1-100 TB but I’d like to make sure we understand the scale of this first.
That’s dangerous dark magic, won’t catch me doing that! I’m only sending good vibes and sunshine from San Diego to you New Yorkers!
…except that isn’t the case. We have backup and off-site backup. It’s third-party backup that’s missing (and isn’t provided by “the cloud”, anyway), so looking into that is potentially worthwhile. Mind, I would really like to see that backup be all of (public) gitlab.kitware, not just VTK. (In particular, I would like to see at least CMake included.)
No need to be nervous. The VTK source code history would be very difficult to destroy completely, as full VTK git repositories are cloned to tens of thousands of users’ computers. Even if entire continents are taken out, VTK would survive. Losing all the issues and pending merge requests and having to reconfigure CI from scratch would be unpleasant but nothing catastrophic.
Thanks for clarifying. That off-site backup either didn’t land with me or it took a while for us to establish that (which is part of the problem I’m outlining around the public committments and transparency).
What I’m looking to see here is whether we can establish an “off-site” backup is actually accessible to the people who’ve contributed to VTK who aren’t employed by or under contract with Kitware.
Also, if you are worried about a situation where Kitware cannot continue maintaining the infrastructure, we make a commitment to work with the community to move the history to a community resource such as Github and Gitlab. So please do not worry about the preservation of this information. It’s covered.
The other face of this is accessibility. That is whether our infrastructure can support all the demand coming from humans and AI agents. Of course, the same applies to Github and Gitlab. It is too early to tell honestly. Let’s give it some time and see where things land. If we all conclude that the current infrastructure cannot reasonably support community development of VTK, we start discussing next steps.
I don’t think anyone is realizing that I’m not in the least bit worried about the source code. I’m worried about the context around the code. The thousands of issues and discussions, many of which we directly link to as explanations for design decisions in PyVista. This isn’t about the code.
For us mere mortals over on the PyVista project, we cannot deduce all of the same feats of engineering that have gone into VTK from source code alone. That’s my point here.
Let me know how big these archives are and I’ll scope this out as well as find the funding to fix this mess and ensure the prosperity of the project for everyone else out there
How much are you willing to preserve? Only VTK? VTK doesn’t exist in a vacuum, nor is it the only project I’d say is deserving of this extra layer of preservation. Can we preserve all of (public) gitlab.kitware? Or at least VTK+CMake+ParaView? (Do we need 2-3 numbers for how much data we’re talking?)
Note, the easiest number to obtain is probably all of gitlab.kitware, including private repositories, which will be a larger number than needed to preserve only the public parts, but probably not more than double.
I am totally in alignment with you that all projects need to be preserved. However, I don’t have a stake in those other projects. To be honest with you, what I am most concerned with is the preservation of VTK specifically and I don’t have the bandwidth or capacity to step up for these other community projects. From my perspective, Kitware should play that role as a leader in open source.
All of these conversations in the other thread and this one have led me to believe that Kitware simply isn’t interested in ensuring that this stuff is reliably available to people outside of Kitware. That’s on you guys to figure out.
For me, and for everyone else who is heavily invested in VTK, I’m willing to figure out and fund something to ensure the preservation of all of that rich history and context. We rely on it as important archival records of decisions and engineering accomplishments that have gone into VTK and support our other downstream work.
If you guys want to have your own initiative that archives all of the projects in the VTK ecosystem, I’m all for it. but frankly, I am concerned and feel a sense of urgency, given how utterly unreliable gitlab.kitware.com has become, that I want to take this into my own hands here to ensure that this is addressed.
Not trying to be rude here, but it isn’t reasonably supporting it today. I have not been able to confidently access gitlab.kitware.com in over 6 months
I’m willing to organize and pursue funding to independently host a publicly readable archive of VTK’s community history.
Could someone at Kitware help me establish:
Size and coverage: Roughly how large would an export of VTK’s public issues, merge requests, reviews, comments, and attachments be? Could we get a separate estimate
for VTK Discourse, and identify any older history still held elsewhere, such as Mantis?
Initial export: Could Kitware provide a public-data-only export to seed the archive, with any known omissions identified? And a timeline that Kitware finds reasonable.
Ongoing access: What authenticated, read-only API access or approved crawler access could be arranged, including an Anubis exemption if needed, to keep it current?
Coordination: Who should email to agree on request rates, concurrency, collection windows, and a way to pause immediately if it affects service?
As far as availability of gitlab.kitware… it’s an ongoing issue. It may be worse externally, but if you think it isn’t affecting us internally, well… that’s incorrect. Nor do I think it’s fair to say “Kitware doesn’t care”.
As far as a third-party mirror… this is a problem that is difficult for Kitware to solve by definition, because anything that relies on Kitware ceases to be “third party”, and thus defeats the purpose.
Regardless, I think the likelihood that the data is going to be lost is acceptably small.
Yeah. That’s why I’m proposing to take the initiative here and backup/mirror the content.
It’s difficult to just take a “trust us, it won’t be lost” with the insight given so far into the state of the the self-hosted infra and supposed offsite backup…
Honestly, my immediate concern is less about the data being lost (though I am concerned about this) and more about the fact that I can’t seem to access it at all the majority of the time. I’m mostly focused on accessibility in this moment.
FWIW this is exactly the sort of statement that puzzles me and should reasonably give pause to outsiders. Why is traffic on an isolated (informational) website making GitLab go down? How porous is the boundary between these systems? Does a rogue WordPress CVE mean GitLab is also compromised?
This seems to bring up the topic of governance and whether multiple organizations (PyVista and Kitware and others) are currently able to provide administration of VTK (or CMake). Is the VTK project currently dependent solely on Kitware or can a non-Kitware organization provide administration of the VTK project? All props to Kitware for all the great work been done, though has VTK grown to be a bigger thing than Kitware itself where multiple orgs could administer the project?
For the 3D Slicer project there’s been experimentation with other CI infrastructure to more easily allow it to be administered by non-Kitware developers. Currently it is reliant on hardware behind the Kitware network which results in the open-source project being dependent on Kitware.
However, PyVista has been volunteer maintained for nearly all of its ~10 year history, so unfortunately it is unlikely our community will have the capacity to provide additional administration of VTK or foster development of VTK itself. This is something others in our community of 100s of contributors and millions of users may be interested in as there are real business interests in keeping PyVista stable, though continuing development of VTK is not something I’m personally that interested in.
My goal is to ensure the posterity of VTK in an archival sense and to make sure that it is reliably accessible to the PyVista community (people, organizations, AI labs training models, agents, etc… non-exclusively).
While my goals/motivations are limited/different, I’m all ears and happy to help in a bigger effort like this any way I can.
But I’d like to come up with a near-term solution to the accessibility issues before focusing on any larger initiatives.