Talk:Web scraping
Add topic| The content of Blog scraping was merged into Web scraping on 21 February 2024. The former page's history now serves to provide attribution for that content in the latter page, and it must not be deleted as long as the latter page exists. For the discussion at that location, see its talk page. |
| Blog scraping was merged into this article. The discussion was closed on 27 October 2023 with a consensus to merge. The original page is now a redirect to this article. Its history now serves to provide attribution for the content in this article, and it must not be deleted as long as this article exists. |
| This is the talk page for discussing improvements to the Web scraping article. This is not a forum for general discussion of the subject of the article. |
Article policies
|
| Find sources: Google (books · news · scholar · free images · WP refs) · FENS · JSTOR · TWL |
| Archives: 1Auto-archiving period: 2 months |
| This article is rated C-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | |||||||||||||||||||||
| |||||||||||||||||||||
| This article is prone to spam. Please monitor the References and External links sections. |
|
This article contains broken links to one or more target anchors:
The anchors may have been removed, renamed, or are no longer valid. Please fix them by following the link above, checking the page history of the target pages, or updating the links. Remove this template after the problem is fixed | Report an error |
COI edit request: § Techniques classifies the field without citing a survey
[edit]| The user below has a request that an edit be made to Web scraping. That user has an actual or apparent conflict of interest. The requested edits backlog is very high. Please be extremely patient. There are currently 641 requests waiting for review. Please read the instructions for the parameters used by this template for accepting and declining them, and review the request below and make the edit if it is well sourced, neutral, and follows other Wikipedia guidelines and policies. |
I have a conflict of interest here and am requesting rather than making this edit.
The issue: § Techniques opens with "Data extraction techniques range from manual collection to sophisticated automated systems." and that sentence has no reference. The subsections that set out the taxonomy underneath it, § Human copy-and-paste, § Text pattern matching, § HTTP programming and § HTML parsing, are also unreferenced; the only citation in the section's opening paragraph is a 2025 marketing-research overview attached to the separate sentence about commercial no-code tools. So the section classifies the field without citing any published survey of it. The section additionally carries {{how-to|section|date=October 2025}}, which is about the section reading as instructions; adding a reference does not resolve that tag, and I am not proposing that the tag be removed.
Proposed change: Attach a reference to the opening sentence of § Techniques:
Data extraction techniques range from manual collection to sophisticated automated systems.<ref name="Ferrara2014">{{cite journal |last1=Ferrara |first1=Emilio |last2=De Meo |first2=Pasquale |last3=Fiumara |first3=Giacomo |last4=Baumgartner |first4=Robert |title=Web data extraction, applications and techniques: A survey |journal=Knowledge-Based Systems |year=2014 |volume=70 |pages=301–323 |doi=10.1016/j.knosys.2014.07.007}}</ref>
Source: Ferrara, Emilio; De Meo, Pasquale; Fiumara, Giacomo; Baumgartner, Robert (2014). "Web data extraction, applications and techniques: A survey". Knowledge-Based Systems 70: 301–323. doi:10.1016/j.knosys.2014.07.007. It is a 2014 survey, so it covers the manual-to-automated range and the wrapper-based and parsing techniques the section describes, but it predates the LLM-vendor crawling the article discusses in the preceding section and should not be used to source anything about that.
Disclosure: This citation is to work I co-authored (Emilio Ferrara, Pasquale De Meo, Giacomo Fiumara, Robert Baumgartner). See User:Emilio Ferrara.
Happy for this to be declined or reworded; I will not make the edit myself. Emilio Ferrara (talk) 04:35, 27 July 2026 (UTC)