Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a23ccaaf0ab02c38

Jump to content

Talk:Web scraping

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia

COI edit request: § Techniques classifies the field without citing a survey

[edit]

I have a conflict of interest here and am requesting rather than making this edit.

The issue: § Techniques opens with "Data extraction techniques range from manual collection to sophisticated automated systems." and that sentence has no reference. The subsections that set out the taxonomy underneath it, § Human copy-and-paste, § Text pattern matching, § HTTP programming and § HTML parsing, are also unreferenced; the only citation in the section's opening paragraph is a 2025 marketing-research overview attached to the separate sentence about commercial no-code tools. So the section classifies the field without citing any published survey of it. The section additionally carries {{how-to|section|date=October 2025}}, which is about the section reading as instructions; adding a reference does not resolve that tag, and I am not proposing that the tag be removed.

Proposed change: Attach a reference to the opening sentence of § Techniques:

Data extraction techniques range from manual collection to sophisticated automated systems.<ref name="Ferrara2014">{{cite journal |last1=Ferrara |first1=Emilio |last2=De Meo |first2=Pasquale |last3=Fiumara |first3=Giacomo |last4=Baumgartner |first4=Robert |title=Web data extraction, applications and techniques: A survey |journal=Knowledge-Based Systems |year=2014 |volume=70 |pages=301–323 |doi=10.1016/j.knosys.2014.07.007}}</ref>

Source: Ferrara, Emilio; De Meo, Pasquale; Fiumara, Giacomo; Baumgartner, Robert (2014). "Web data extraction, applications and techniques: A survey". Knowledge-Based Systems 70: 301–323. doi:10.1016/j.knosys.2014.07.007. It is a 2014 survey, so it covers the manual-to-automated range and the wrapper-based and parsing techniques the section describes, but it predates the LLM-vendor crawling the article discusses in the preceding section and should not be used to source anything about that.

Disclosure: This citation is to work I co-authored (Emilio Ferrara, Pasquale De Meo, Giacomo Fiumara, Robert Baumgartner). See User:Emilio Ferrara.

Happy for this to be declined or reworded; I will not make the edit myself. Emilio Ferrara (talk) 04:35, 27 July 2026 (UTC)Reply