Skip to main content
kfitz

In the Wake

Crossposted from the Knowledge Commons team blog.

Yesterday, we published a message to the Commons community explaining our recent downtime. In the interest of transparency, I want to dig a little further into what happened, how we responded, and what we have learned. The entire Commons team welcomes your feedback.

What Happened?

On the afternoon of October 1, 2026, an attacker began scanning our WordPress installation for security vulnerabilities. Unfortunately, there were vulnerabilities to be found; we have known for some time that a few of our dependencies (various pieces of lower-level software in our stack) had security issues, and we have been working over the last year to close those holes. Those issues could not be resolved by simply bumping the software to its next version, because every new version of a small piece of software brings with it the potential for breaking something unexpected in a complex network like ours. Not only are we operating with a complex codebase, but our users are building their sites on top of our platform, and updates run the risk of breaking what they have done. As a result, updating our codebase requires substantial effort to ensure that the upgrades are done safely, as well as to test the results and remediate any problems. That work was almost complete, but we ran out of time.

As soon as we saw the first evidence that something was happening, we began investigating to see if we could determine whether we were experiencing a problem with our codebase or an attack. In the early morning hours of October 2, eastern time, we became suspicious that we were experiencing a malicious intrusion, and shortly thereafter we discovered clear evidence that we were correct. Martin Eve, our associate director for technical development, ordered an immediate shutdown of the platform, a rotation of all of our access keys, and a revocation of our CILogon authorization keys so that no personally identifiable information could be exfiltrated from that platform. We then sent an update to our governance council, notifying them that we had experienced a data security incident.

Over the ensuing days, we formulated a secure path to recovery and worked around the clock to implement it. After conducting a full audit of KC Works and KC Profiles, we were confident that those systems had not been accessed and that they were secure and ready to be brought back online, which we accomplished on October 5.

Having localized the attack site to our WordPress instance, we audited our access logs and examined other aspects of our server records to determine whether data from that system had been exfiltrated or otherwise affected. We found no evidence to that end. We then set about ensuring that all our dependencies, plugins, and themes were fully updated, and that none posed additional vulnerabilities. A few plugins and themes had to be removed from our system, as their maintainers had not remediated known security gaps.

We were concerned, however, that the attacker could have injected malicious code into our system without it being readily spotted; such injections can be particularly dangerous because they can lie dormant for periods of time, or even replicate themselves throughout a database or codebase, before being activated. To ensure that the database was clean, we rolled it back to the last backup we had taken prior to the first evidence of the intruder's presence. As a result, we may have lost up to 12 hours' worth of WordPress activity, for which we apologize. However, no changes made to Works or Profiles during that period were lost.

On the morning of October 8, the network was fully relaunched and communication with our userbase about what had taken place began, including working with users whose sites or groups have been affected by the code upgrades.

What Next?

Having completed cleanup and relaunch, we're now looking at all of our security provisions to ensure that we are as hardened against future attacks as we can be. This includes building a feed of our dependencies and any security updates that might apply to our stack, establishing a daily practice of reviewing and, where appropriate, updating to patch any known vulnerabilities, and hardening all of our security practices and apparatuses as much as we can. The work of the last week puts us in a much better position to respond to new updates more quickly, which will be a big help. However, we can’t simply upgrade all of our dependencies as soon as new releases come out, given the potential for supply chain attacks, in which malware can be smuggled into a new release, so we have to examine those releases carefully before implementing them. It's necessary work, but for a team as small as ours it's time-consuming, and it will undoubtedly slow our progress on other project goals.

But even so, we are still left with questions, many of them driven by the technical environment in which we all live today. For instance:

Was the attacker human? A human assisted by AI? Or an AI bot "gone rogue"?

I will start by noting that I hate this idea that a bot can "go rogue"; as Martin said to me the other day, echoing similar statements by Cory Doctorow, having a bot escape containment does not mean that it has willfully sought escape but rather that its developers have set up crappy containment. Or, as Mastodon user jwz@mastodon.social recently posted, "You park your car at the top of a hill. You leave a note on the dashboard that says 'please be good' and then you pull the parking brake. The car careens down the hill, smashes everything up, and you put out a press release saying 'My car ignored my instructions! It hallucinated! It went rogue!' Your insurance company not only believes you, but invests $100 million in your company."

That having been said, it is clear that the attack on Knowledge Commons was at minimum AI-assisted. The attacker scanned our codebase and was able to chain together several small unrelated weaknesses into a large-scale vulnerability with a speed that no human hacker working alone could manage. As Martin notes:

One of the things that has changed as well here is that vulnerabilities are exploited within hours of their release. There's no heads-up time for us to spend ages figuring out whether we should do this upgrade and whether it breaks things anymore.

It used to be the case that when the vulnerability was declared, hackers then had to spend ages reading the code that showed what the vulnerability was, then determining how to exploit it, what that route would be. And that gave the people running the software enough time to evaluate it, integrate it and do the upgrade.

That doesn't exist anymore because the LLM can work out what the vulnerability was and how to exploit it within a very short space of time.

It's a brave new world, with a lot of potentially malevolent forces including but not exclusively people in it.

How much of this attack is attributable to Knowledge Commons being an open-source project?

Kind of a lot? Having all of our code out there in the open is what makes it scannable. But our ability to recover is also attributable to our open-source stance. On the one hand, the attacker (human and/or bot) did have access to our codebase, which was right out there to be scanned. On the other hand, our reliance on open-source software means that there are many good-actor eyes on the code as well, spotting and patching vulnerabilities as they occur.

We are committed to open-source systems and methods, not just because of our commons-oriented values, but also because we believe that these systems and methods produce better software. Just as Lots of Copies Keep Stuff Safe in digital preservation, lots of eyes make more bugs (and vulnerabilities) shallow. However, open-source communities are facing a deep reckoning right now regarding how their work can go forward as openly as possible while also being as secure as possible.

We have recently submitted a proposal to the NSF's PESOSE program -- or Pathways to Enable Secure Open-Source Ecosystems -- in order to develop a true open-source development community around Knowledge Commons. We hope to have the opportunity to work directly with members of our community to think about how the process of building and maintaining the Commons can be both open and secure moving forward.

What else?

What further questions do you have? Let us know in the comments; we'd very much like to keep a discussion of these issues going.

And again, we thank you for your support, and we apologize for the inconvenience that this hack and the resulting process of remediation created. We continue to focus on doing better for all of our community members.

Webmentions

No replies yet.