We use affiliate links. They let us sustain ourselves at no cost to you.

OxyCon 2026: A Recap

Our impressions from Oxylabs’ seventh annual conference on web scraping.

Adam Dubois

OxyCon 2026 has ended; as always, we’re here to cover it. This year’s format stayed close to its roots – but the real draw, of course, lies in the content. To spoil things for you a bit, we’re moving away from AI maximalism to more level-headed considerations about its costs and capabilities. 

This article will recap all seven presentations and two panel discussions, so you can get acquainted before watching the full recordings.

You'll find our coverage of earlier OxyCons and other major industry events here.

Organizational Matters

Oxylabs doesn’t like reinventing the wheel. The conference was filmed on location and had a stage; but like most years before, most attendees experienced it online. It took place on September 16 in a time zone catering to European audiences. 

After registering for free, receiving an invitation link, and logging in, you had a video stream with a Q&A box below it. Questions could be upvoted, but they required approval to appear, which took some fun out of it. Overall, the event proceeded smoothly.

oxycon 2025 platform
The platform for online viewers. We took this image from last year's recap, but it hasn't really changed.

Main Themes

Quite inevitably, the narrative was dominated by artificial intelligence. We saw it power a web scraping tool, help a company’s customers, as well as simplify internal engineering processes. It came up in legal and even philosophical contexts. 

But while large language models continue hogging the mindspace, we feel like the attitude towards them has changed. It’s become less cavalier, focusing less on how we can use the most of this incredible technology right now in any way we can and not fall behind; instead, people have started raising concerns about the constraints, costs, and the relationship between the person and the AI. 

In addition, websites continue accelerating their efforts to protect valuable data, with Google – one of the largest sources of them all – now taking the helm. It creates serious threats but also big opportunities. However, the conference had very little to say about search indexes trying to usher in the next stage of AI adoption while at the same time upending the traditional models of content distribution. 

Personally, we prefer this more level-headed approach, no matter how amusing the articles about the hangovers of tokenmaxxing may be.

The Talks

Here are this year’s presentations and panel discussions. Select a title to jump there. 

Talk 1: Your Next Dataset Starts with a Prompt

A product preview if we’ve ever seen one. Andrius Kuksta from Oxylabs demonstrated how to scrape and parse 1,000 products using Oxylabs’ infrastructure, AI skills, and a user interface layer mounted on top. This way, engineers can onboard a new website in half an hour rather than a week (Andrius described the process in a rather witty leadup). 

The tool in question started off by taking a URL, which can be a homepage. Step one involved asking some questions to generate a navigation plan. Step two was writing the code using a file system that resembled an IDE. Step three had it build a parser using a selection of URLs; they served as a training dataset. Step four culminated with running the pipeline, which can be done either locally or on hosted infra. Each step had its own underlying skills trained for the best effect. 

Apparently, the dataset builder (we didn’t catch the name) is already available to use, but you need to contact an account manager. Watch this if you’re interested in using such a tool – or if you’re building something similar.

oxycon 2026 talk 1
A back and forth with the AI, only without the forth appearing in the chat window.

Talk 2: Web Hydration: Loading Dynamic Content without a Browser

In this talk, Etienne Ellie, Lead Developer at ScrapingBee, tried to persuade viewers to use headless browsers more judiciously. Instead, he proposed turning to JavaScript engines which suffice surprisingly often and are much less expensive to run.  

Though Etienne focused on JSDom as a way to hydrate pages, the information he presented can be generalized to all similar tools. He explained how JS renderers work, gave concrete examples of their benefits (less memory, more concurrency, baby!) and limitations (harder to forge fingerprints). The main takeaway invites you to climb a complexity ladder, reaching for browsers last. Worth a watch.

oxycon 2026 talk 2
JavaScript engines JSDominate full browsers in resource use.

Talk 3: Where AI Actually Helps in Production Scraping

Giedrius Steimantas from Oxylabs took the role of a disgruntled engineering manager who’s had one handwavy “just use AI” suggestion too many. As tokenmaxxing goes out of fashion, throwing everything into AI can lead to very uncomfortable talks with finance when you’re parsing 700M pages daily. 

The reality was, of course, much less dramatic. Giedrius recounted a failed internal vibe parser tool which was supposed to speed up data structuring but ended up requiring extra engineering effort and fell into disuse. Having learned their lessons, his team then built a vibe unlocker which combined a CLI, tribal knowledge put into skills, and strict guardrails. Unlike the first try, the vibe unlocker turned out to be a resounding success. Giedrius explains the reasons – and his future plans – in more detail.

oxycon 2026 talk 3
Happy wife engineers = happy life company.

Talk 4: The Cost of Solving the Wrong Problem: Data-Driven Decisions for Web Access

Juan Manuel Perez and Kieron Spearing from Centric Software talked less about the consequences of tackling the wrong problems, and more about identifying and solving the real needle movers – so basically, the other way around. Their problem space was anti-bots, at least in the context of the presentation. 

In essence, the duo walked viewers through its tango with Akamai, going back to 2022. This anti-bot system consumed a half of their feeds, over 50% of the proxy cost, and made the crawlers break often. Juan and Kieron described the pivotal moments, particularly when Akamai broke thousands of their crawlers in 2025. They shared their broadly applicable troubleshooting steps and the main learnings for successfully scraping in 2026. 

In this case, the QA was as insightful as the presentation itself, as the presenters explained why they built solutions in-house (it’s about control), why geolocation coherence is key, and their biggest current challenges.

oxycon 2026 talk 4
Uh-oh. And we’re not referring to anthropomorphized AI dogs.

Talk 5: Agentic Workflows: Patterns Behind Successful AI Automation

Intentionally or not, Yenny Chueng’s (Bluefish AI) presentation captures the zeitgeist of AI use really well. She gives the engineering perspective of a company that can’t afford to run fast and break things – that’s because Bluefish caters to compliance-requiring F500 companies (around 10% of them). 

Yenny’s talk is neatly structured, so we’ll give you the scaffolding and let the recording fill in the gaps. Despite all the commotion about AI agency, humans remain at the center of decision making, owning areas that require planning, creativity, and most importantly – taste. The primary metric to use shouldn’t be usage but rather cost, while the most important thing to optimize for is the outcome. Finally, skills serve as a great medium for standardizing and transferring earned knowledge.

Personally, this presentation is a must – at least it was for us.

oxycon 2026 talk 5
Humans still have it in them.

Panel 1: The Adaptive Web: Turning Industry Challenges into Competitive Advantage

Pierluigi Vinciguerra of The Web Scraping Club fame, Hocine Amrane from NielsenIQ, Andreea Stroe from Adobe and Andrey Gourine from Limy AI sat down to talk about the status quo from an engineer’s point of view. The discussion was moderated by Marica Gecaite, Chief Commercial Officer at Oxylabs. 

The discussion kicked off with bots and the agentic web. Even though the former already outnumber human traffic, the latter remains a buzzword. The participants then turned to the topic of scraping LLM output, which introduces unique conundrums, such as non-deterministic answers and data poisoning. 

Another question that inevitably arose involved using LLMs in workflows. For now, they’re best at generating boilerplate code for the multitude of unprotected websites, and many problems can still be solved by proper automation. The role of AI in professional training also received attention: Pierluigi cautioned against relying on AI too much early on. 

When asked to finish a sentence about the future of scraping, most participants answered that scraping is going to get harder and more expensive. Everything else remains to be seen, and you should see this panel discussion.

oxycon 2026 panel 1
The panelists.

Talk 6: We Made Our Web Search Worse on Purpose

Emilis Strimaitis, Head of Innovation at Hostinger, embodied the sin of gluttony. His was the tale of stuffing customer-facing AI tools with everything, learning lessons, and gradually removing the cruft to regain sanity.  

Emilis’ journey was eventful. Throughout it, his team learned not to overstuff the context window and that no one needs to remember everything. Absorbing those lessons reduced token use and improved results. They also had a whopping 160 tools that the agent would check each time, together with skills for every minor action. If that wasn’t enough, Hostinger had made seven specialist helper agents that would send customers on a Kafkaesque journey if their functions overlapped. Oh my. 

In the end, though, temperance prevailed, and everyone is happier for it. If you’re building something with AI (yes, you are), watch this to laugh and to learn at someone else’s expense.

oxycon 2026 talk 6
No caption is necessary.

Panel 2: Is the Web Closing? What It Could Mean for Businesses, Journalists, and Researchers

The web is obviously becoming harder to scrape. Many of these issues relate to tech; do they also carry over to policy? That’s what Carl Miller, Fellow and Co-Founder of the Centre for the Analysis of Social Media (that’s a mouthful), Paul Bradshaw, journalist and academic at BBC & Birmingham City University, Krishna Sood from the European AI lab Black Forest Labs, and Oxylabs’ own Denas Grybauskas gathered to discuss. 

Denas kicked off by demonstrating the growing number of restrictions in robots.txt files and terms of service, then stating that copyright-related lawsuits have tripled in the last year to nearly 150. He then set the tone with a question: is the web really closing?

As the job titles suggest, the panel brought highly diverse points of view and their own concerns to the table. For Krishna, data access was becoming a big challenge, with data becoming gated by log-ins and large platforms using data but not sharing it with others. Licensing doesn’t really help due to financial asymmetry, and there are strong public policy reasons for keeping data accessible to everyone. 

As a journalist, Carl lamented about finally getting the tools for analyzing large corpora of data, but having fewer ways to get in the first place. The irony is that 15 years ago it used to be the other way around. And even though there’s a European law that allows requesting data from social media platforms, they engage in full-on malicious compliance and stalling tactics. Carl treats robust data access and analysis as crucial tools for maintaining democracy. 

Paul actually had a different opinion: to him, data has never been open, so now academia finds it easier to get thanks to AI. But he too had reservations, especially concerning licensing deals. In the end, Paul wasn’t convinced that scraping someone gives others blanket permission to scrape you. 

We always recommend panel discussions. This one has received our longest write-up, showing that we consider it important and relevant for all involved in our niche.

oxycon 2026 panel 2
The other panelists.

Talk 7: Best Open-Source Scraping Repositories of 2026

The final presentation was right up our alley – it was a research-based listicle! Tadas Gedgaudas, a former Oxylabs employee and now a business owner, compared nine scraping repositories by their ability to open protected websites. 

Tadas divided the tools into three categories: HTTP scraper engines like Rnet, full browser engines like Camoufox, and agent scraper engines, such as Lightpanda. The naming obfuscates it, but the last category basically included JavaScript rendering engines that weren’t full browsers. Tadas ran each of these tools against 100 websites with and without proxies, measuring three metrics: success rate, speed, and resource use. 

The presenter then introduced an open-source benchmark called ScrapingArena. All in all, we found the presentation interesting, but the format and time constraints made it too lossy to really explain why the tools performed the way they did or to make a serious comparison.

oxycon 2026 panel 7
ScrapingArena may not give you all the answers, but it can put you on the right path.

Conclusion

That was all for this great conference! 2026 has one more major event in store, Zyte’s Extract Summit. Its two chapters take place on October 7-8 in Austin and November 10-11 in Dublin.