Adversary spaces are closing. The most useful online environments are becoming harder to observe through routine public collection, for several reasons. Previously, we’ve written about the impact of Australia’s Under 16s Social Media Ban. Now, eight months later, other countries, including the UK, are following suit, and the difficulties in maintaining access to platforms is a key concern for law enforcement and OSINT practitioners. Additionally, the high profile of open-source collection means that bad actors are increasingly aware of the risks of public platforms. OPSEC isn’t just for the white hats, after all.

Mounting Challenges for OSINT Practitioners

When actors move conversations into encrypted messaging services and private channels, information that was previously in the public sphere is no longer accessible. In other cases, a social media space is still reachable, but only with the right tradecraft. Discord is a good example. Valuable content may sit in servers that are searchable and rich in context, but access often requires an account, sometimes participation, and a good deal more OPSEC planning (check out our Discord blog here).

Another challenge for OSINT practitioners, particularly those working in the national security and corporate space, is navigating country-specific ecosystems such as the Chinese internet. Here, the challenge is less about pure access and more about platform familiarity, language, context and safe collection (though access is still an issue, with many region-specific sites geo blocked for users outside mainland China).

Meanwhile, the internet continues to grow. As of this year, 73.8 percent of the global population uses the internet. In 2022, it was 62.5 percent – in four years, that’s more than a billion extra people. The firehose of public information still exists, but it’s scattered across an increasingly complex online environment.

What do we do when we know the data is there, but we don’t know how to retrieve or analyse it? It feels like a contradiction – we face a challenge of having too much information, while closing spaces, at the same time, limit our visibility of online content. In this blog, we’ll look at a multi-tool approach and discuss some of the evolving challenges for data collection.

The Multi-Tool Approach

A willingness to use multiple tools, each suited to a specific task within the OSINT workflow, is the first step to improving our collection capability. This approach can comprise some or all of the following:

  • Manual searching for leads generation and pivot points. Being able to effectively navigate the online environment and make judgements is more important than ever.
  • Easy scraping tools and extensions: use these to collect representative data sets, but to keep things low-cost, expect to come away with limited datasets.
  • AI collection workflows – internet connected AI can retrieve representative data from platforms of interest, and navigate regional geo-blocks.
  • For niche or hard-to-access data, bespoke collection tools or curated data sets may be the best option. Importantly, though, we need a plan for what we do with the data when we procure it.

Manual Search and Pivoting

While tools can enable collection at scale, manual searching techniques remain key when it comes to locating relevant content. Manual searching can, of course, be combined with AI research and web retrieval tools, but verification continues to be an issue, here. Below is a screenshot of information retrieved about a company website (albeit a fake company website). While AI did a sound job of grabbing contact information and key points, it confidently and inaccurately listed a location (Varanasi, in India) that doesn’t appear anywhere on the site itself.

Model Google Gemma 4, 4 billion parameters running on Ollama, interface is OpenWebUI (formerly Ollama Web UI)

Building out keyword lists informed by subject matter expertise and leveraging advanced search skills (Google dorking and equivalents across multiple search engines), as well as incorporating verification (double check ALL sources) into manual workflows means that you’re finding the right platforms and the right pivot points.

Low Effort Scraping Tools

Scraping is sometimes touted as the OSINT panacea: when it comes to capability requirements, ‘scraping’ is often at the top of the list. Social media scraping, link scraping, search scraping – this approach allows for data collection at scale, and more data (sometimes!) offers more insights. Five years ago, free, targeted scrapers for specific platforms were available as command-line tools, although often required technical know-how. Twint, for example, was a Twitter scraper that required minimal set-up, and could efficiently retrieve hundreds of Twitter posts based on queries within seconds. It didn’t require users to be logged in, which was a boon for OSINT practitioners with restrictions around account registration.

Anti-scraping measures and API restrictions have relegated Twint (and other third party tools) to the scrapheap; additionally, many practitioners working on corporate systems may encounter blocks when setting up virtual environments for large scale collection. Social media scraping is still possible, though – and with the arrival of AI, entry level tools are widely available and require very little technical knowledge.

Scraping Extensions:

  • Simplescraper: https://simplescraper.io/
    • Handy for social media comments and link extraction, and no specific skills required.
  • Web Scraper: https://webscraper.io/
    • As above, although slightly more complex interface – requires a bit more upskilling
  • Thunderbit AI-enabled scraper: https://thunderbit.com/
    • Though free collect is limited, the pre-built recipes and AI integration make this an excellent choice for gathering sample comments from YouTube videos and some social media platforms.
Scraped YouTube Comments from an Explosive Media pro-Iran video.

AI Actors for Scraping

Apify (https://apify.com/) is the standout here for a few reasons. It’s framed as a one-stop shop for AI scraping tools, and for good reason. It’s a marketplace for AI ‘actors’ – agents that retrieve the data you’re after. Tools are run through the Apify console, and an account is required. There is a free tier, though this requires registration (i.e. a Google account) as well as paid plans. Depending on collection requirements, the free tier might be all that’s needed – it’s particularly useful for one-off collection tasks or person of interest investigations (when entities have public profiles). Below is an example of scraped Facebook posts from a public page, with post URLs and metadata fields.

AI-enabled scraping allows easy, fast collection from a range of platforms, without the learning curve; however, full result sets will cost. Anything beyond a few hundred results (using reputable, customised scraping tools, at least) will come with a price tag.

AI-Enabled Direct Collection

Though many popular scraping extensions leverage AI, we can also use internet-connected LLMs to collect directly from platforms of interest – the key is to choose the right model. Earlier this year we published a blog on Reddit tradecraft and included examples of quick and easy post collection based on keywords. We can perform similar collection for platforms like Facebook and X, though levels of success can depend on the interface used.

X (Twitter) Collection Using Grok

Despite the ups and downs of the past few years, X continues to be a valuable source of public data when it comes to news and trends. It’s particularly relevant for anyone researching disinformation and scams or monitoring conflicts.

For basic X collection, Grok is the standout option. While bulk data collection is still off limits (Grok cannot bypass Twitter’s API rate limiting), it can perform batch collection. For example, we can retrieve posts based on date filters (to avoid overlap) in batches of ten to build a representative sample data sets and develop leads.

Region Specific Collection

For OSINT practitioners conducting research in regions with access restrictions (China, Russia and Iran are the key regions, here, although any region-specific search requiring native language collection presents challenges), choosing the right retrieval tools can help to navigate geo-blocked content.

China’s internet is the most notorious when it comes to access issues. Most people are familiar with the concept of the Great Firewall, which heavily censors content and access for users within China. Users outside China, by contrast, are likely to encounter the Reverse Great Firewall when attempting to retrieve content from Chinese websites and databases. A recent study on the Reverse Great Firewall found that a number of methods were used to block external access to Chinese government websites, including explicit geo-blocking, timeouts, and failed DNS resolution. Additionally, most Chinese government, academic and corporate sites have a Chinese version, designed for Chinese audiences and often inaccessible outside of China, and a version for overseas users with drastically reduced data available.

This is where AI content retrieval can assist. While Western AI models are perfectly adequate for some China-specific collection and generally have good (if not nuanced) translation capabilities, leveraging Chinese AI models assists practitioners with collection from Chinese, rather than Western, sources.

Though not all Chinese AI models are accessible outside China, most of them are (or at least a version of them). Below is a list of the top models, along with access details (as of July 2026).

Open-weight status and account-free access from Australia — as of July 2026. Sources: official model cards/Hugging Face licenses, vendor sites, and press coverage as of July 2026. “No-account access” reflects publicly available guest/no-login modes. Regional access restrictions may change.

Bespoke Tools and Curated Data

Would you rather pay small sums of money to different platforms for scraped results, or pay a larger sum of money to have it all in one place? The answer is probably dependent on factors like budget, size of your team, how consistently you need to scrape data, security requirements, etc.

Finally, there’s the question of bespoke tools that provide multi-platform collection, ongoing monitoring, analytical tools, and commercial datasets. Tools like NexusXplore streamline collection and analysis and provide access to data that is difficult or impossible to retrieve through manual tradecraft.

Features of bespoke tools may include:

  • API access for multiple platforms means that one tool can collect social media posts for a variety of platforms – no need to switch between extensions or LLMs.
  • Monitoring capabilities for ongoing collection – automate regular collection workflows to stay on top of emerging information.
  • Operational security and attribution management (this is particularly relevant to government and law enforcement teams where risk of attribution is a significant blocker to open-source collection).
  • Analytical tools including data visualisation and AI integration for triaging collect.
  • Curated and commercial datasets that aren’t available on the open web.

While teams may not use all approaches for every investigation, leveraging a range of tools for collection (and knowing which ones work best for the type of information required!) is vital for practitioners working in an increasingly complex online environment.

If you’d like to take your OSINT collection and tradecraft to the next level, get in touch with us at [email protected] or explore our training courses to find the perfect fit for your organisation.