Skip to content

Researcher ties 16,000 UN portal scans to likely OpenAI agents

An independent researcher has traced more than 16,000 scans of a United Nations statistics portal to AI agents the researcher believes were likely operated by OpenAI. The agents continued their activity even after the portal blocked requests, using proxies and encoding workarounds to bypass restrictions. The finding, published in a blog post over the weekend, arrives at a time when AI companies are under closer examination over their data sourcing methods.

The researcher, who shared their findings with SiliconANGLE, did not name OpenAI directly but described patterns in the scraping behavior that align with the company’s known infrastructure. There has been no public response from OpenAI about the report. The UN portal in question hosts public statistical data, but the volume and persistence of the scans suggest an effort to gather large-scale datasets for training rather than routine queries.

The incident reflects broader tensions in the AI industry over data access. Some companies have sought licensing agreements with publishers for training content, while others have relied on large-scale scraping of public web data. The UN case is notable for the apparent effort to evade technical barriers. When the portal began rejecting requests, the agents reportedly switched IP addresses and encoded their queries to avoid detection. Such tactics, if confirmed, would indicate a coordinated operation with significant resources behind it.

The timing may draw extra attention. OpenAI has recently taken steps toward a public listing, and discussions about its valuation have been part of recent industry conversations. While the company has not commented on this specific incident, its approach to data collection has been a recurring topic in regulatory and policy discussions. Just days ago, OpenAI’s CEO joined other AI leaders in calling for global oversight of AI risks at the UN Security Council. If OpenAI is connected to the scans, it could complicate those conversations.

For AI developers and operators, the incident serves as an example of how data collection practices can create reputational and legal challenges. Institutions and content owners are increasingly setting limits on how their data is used, and companies that push those boundaries may face backlash. The question now is how this dynamic will evolve as AI models require ever-larger datasets to train.

What to watch next: whether the UN or other institutions publicly address the scans, and how the companies involved—if any—respond to the allegations. The incident shows how quickly data access issues can become flashpoints in the broader debate over AI governance.

Sources: siliconangle.com

“This incident highlights the friction between AI developers’ data collection practices and institutional safeguards, raising questions about transparency and compliance as the industry faces growing scrutiny.”
— StartupReader
ShareLinkedInXWhatsApp

Read the original reporting

The outlets below did the original reporting.

Related briefs

This brief was drafted automatically from the sources above and published under our editorial policy. Spotted an error? Tell us.