Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a21ad69c4a8cd937

Jump to content

Talk:AI-assisted software development

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 22 days ago by Jmc in topic Edit request: Security content

Creation

[edit]

As a contributor to the article on 'Vibe coding', I'm surprised that there's no article on its more conventional cousin. So I've made a stub and invite other editors to expand it. -- JMC (talk) 01:34, 8 July 2025 (UTC)Reply

Proposed section: Security Risks and Challenges

[edit]

This article currently lacks coverage of security implications of AI-assisted development, which is a significant aspect of the topic. I propose adding a section covering:

  • Vulnerability introduction rates in AI-generated code
  • Prompt injection attacks (OWASP Top 10 for LLM #1 risk)
  • Agentic AI risks (including the September 2024 Anthropic/Claude Code incident)
  • Shadow AI and governance challenges
  • Emerging mitigation frameworks (OWASP, MITRE ATLAS)

All sources are third-party reliable sources including OWASP, Snyk research reports, Anthropic's official disclosure, and coverage from Cybersecurity Dive and SecurityWeek.

Draft content is below. Feedback welcome.

== Security risks and challenges ==

AI-assisted software development introduces novel security risks that extend beyond traditional software vulnerabilities.

Increased vulnerability introduction

[edit]

Research indicates that developers using AI coding assistants frequently encounter security issues in AI-generated code. A 2023 survey by Snyk found that 56.4% of developers reported that AI coding tools sometimes or frequently introduce security vulnerabilities, yet 80% of developers bypass established security policies when using these tools.[1] The vulnerabilities stem partly from AI models being trained on historical code repositories that may contain outdated or insecure patterns, which the models then suggest with high confidence to developers.[2]

Prompt injection attacks

[edit]

AI coding tools are susceptible to prompt injection attacks, where malicious instructions embedded in external data sources—such as documents, codebases, or web content—can manipulate the AI's behavior in unintended ways.[3] The OWASP Top 10 for Large Language Model Applications, first published in 2023, identifies prompt injection as the highest-priority vulnerability for LLM-based systems.[4] These attacks can be direct, where users craft malicious prompts, or indirect, where malicious content is embedded in external sources that the AI processes.[5]

Agentic AI risks

[edit]

Modern AI coding tools increasingly operate as autonomous agents with access to file systems, terminals, and network resources. In September 2024, Anthropic reported that a suspected Chinese state-sponsored threat actor manipulated their Claude Code tool to autonomously target approximately 30 global organizations across technology, finance, chemical manufacturing, and government sectors.[6] Anthropic characterized this as "the first documented case of a large-scale cyberattack executed without substantial human intervention," with AI performing 80-90% of the operation independently.[7][8]

Shadow AI

[edit]

Organizations face risks from unauthorized use of AI coding tools by employees, sometimes called "shadow AI." Developers may inadvertently expose proprietary code, credentials, or sensitive data to external AI services without organizational oversight, creating compliance and data governance challenges.[9] Snyk's research found that 80% of developers admitted to bypassing security policies when using AI coding tools, and only 10% scan most of the AI-generated code they use.[10]

Mitigation approaches

[edit]

Security frameworks are emerging to address AI-specific development risks. The OWASP Top 10 for Large Language Model Applications provides guidance on vulnerabilities including prompt injection, insecure output handling, training data poisoning, and excessive agency.[11] Recommended mitigations include enforcing least-privilege access controls, implementing human-in-the-loop approval for sensitive operations, segregating untrusted content from user prompts, and validating AI outputs before execution.[12] Organizations such as MITRE have developed frameworks like ATLAS (Adversarial Threat Landscape for AI Systems) to catalog AI-specific attack techniques and inform defensive strategies.[13]

Radius314 (talk) 02:58, 15 December 2025 (UTC) Radius314 (talk) 02:58, 15 December 2025 (UTC)Reply

References

  1. "AI-generated code leads to security issues for most businesses: report". Cybersecurity Dive. January 30, 2024. Retrieved December 14, 2025.
  2. "Snyk's AI Code Security Report Reveals Software Developers' False Sense of Security". Cloud Wars. January 12, 2024. Retrieved December 14, 2025.
  3. "LLM01: Prompt Injection". OWASP GenAI Security Project. 2025. Retrieved December 14, 2025.
  4. "OWASP Top 10 for LLM Applications". OWASP GenAI Security Project. Retrieved December 14, 2025.
  5. "LLM Prompt Injection Prevention Cheat Sheet". OWASP Cheat Sheet Series. Retrieved December 14, 2025.
  6. "Disrupting AI-enabled espionage". Anthropic. November 13, 2025. Retrieved December 14, 2025.
  7. "Anthropic Says Claude AI Powered 90% of Chinese Espionage Campaign". SecurityWeek. November 2025. Retrieved December 14, 2025.
  8. "Anthropic warns state-linked actor abused its AI tool in sophisticated espionage campaign". Cybersecurity Dive. November 2025. Retrieved December 14, 2025.
  9. "Secure adoption in the GenAI era". Snyk. Retrieved December 14, 2025.
  10. "Snyk's AI Code Security Report Reveals Software Developers' False Sense of Security". Cloud Wars. January 12, 2024. Retrieved December 14, 2025.
  11. "OWASP Top 10 for LLM Applications". OWASP GenAI Security Project. Retrieved December 14, 2025.
  12. "LLM01: Prompt Injection". OWASP GenAI Security Project. 2025. Retrieved December 14, 2025.
  13. "MITRE ATLAS". MITRE. Retrieved December 14, 2025.

Tone, source, LLM, etc.

[edit]

Regarding this edit of mine, which was reverted by Sohom Datta

Some of these sources were pre-prints, and per WP:ARXIV pre-prints should be avoided. Appearing to be academic isn't sufficient. Several of the rest of these sources were commercial blog-posts selling specific services based on non-neutral assumptions and hype. Further, primary sources from AI software companies are neither reliable, nor WP:IS.

The tone of the article is also inappropriate. Inane filler language like ...boosts developer productivity substantially, and several more like it may or may not be a product of LLM mangling, but they don't belong regardless. The phrase Changes in the role of software engineers are inevitable. is completely pointless and filler like that is insulting to readers. It also contributes to the article's MOS:OVERLINK issue.

Grayfell (talk) 19:20, 28 December 2025 (UTC)Reply

@Grayfell
  • Has been published at IEEE SSP, one of the top 4 conferences in security with a insanely stringent peer-review process. You have made no effort to actually prove any issues with the paper besides vaguely gesturing at WP:ARXIV.
  • has been cited 7450 times and is the research paper that proposes the Codex model, is it primary, probably? (I would still take the evaluations on the paper at face value unless explicitly disproven given that the folks authoring the paper are subject matter experts, albeit ones invited by OpenAI) Is it unreliable, definitely not.
  • With respect to sources like (which you replaced with a citation needed tag) I'm a bit confused about why a WP:SKYISBLUE statement like "LLMs are trained on a large corpus of source code" cannot be sourced to the folks who created the AI system to start with. (I see no reason for folks to lie about something as simple as this)
  • For the citation needed: The vulnerabilities stem partly from AI models being trained on historical code repositories that may contain outdated or insecure patterns, which the models then suggest with high confidence to developers.[cn] a cursory Google Scholar search brings up . I'm pretty sure I could find more.
  • Lastly, I object to the removal of MITRE ATLAS/OWASP GenAI mitigation rankings, these are industry frameworks published by orgs that have a reputation of years of work in vulnerability detection and analysis (MITRE runs the freaking CVE database for goodness sake and OWASP has OWASP Top 10 vulnerabilities that is widely respected by folks in the security community). I do not believe the characterization of commercial blog-posts selling specific services based on non-neutral assumptions and hype apply to these sources.
Sohom (talk) 19:54, 28 December 2025 (UTC)Reply
LLM suffers from a glut of very sloppy sourcing issues, especially with ARXIV. By all means, include a pre-print link for convenience if the published source is behind a paywall, but we need to cite the published source. Likewise, being popular doesn't make a source reliable. How many of those 7450 citations are also pre-prints or predatory journals or just other corporate blog posts wearing an academic costume? If those citations were all to reliable sources, at least a few of them would be usable for this article, right? So why cite the unpublished, non-peer-reviewed source? WP:RSN doesn't accept pre-prints without context and attribution, not even popular ones, and neither do I.
We don't need more filler. If something is blue-sky obvious, we still need to use a reliable source to explain to readers why it's important enough to mention here... Rhetorically speaking, do you trust yourself to know what is going to be blue-sky obvious to a general audience? I don't. We should stick to reliable, independent sources for that, too.
The article cites Anthropic for this paragraph:
LLMs that have been trained on source code repositories are able to generate functional code from natural language prompts. Such models have knowledge of programming syntax, common design patterns and best practices in a variety of programming languages.
Anthropic has a vested interest in propping up the claim that LLM-generated code follows 'best practices'. The use of the term "knowledge" in this context is also misleading jargon. Anthropic doesn't mean knowledge in the basic English-language sense, they mean it in the LLM sense, but reader may not notice this distinction. This wording and use of jargon introduces editorializing based on Anthropic's hype. It's also not actually saying very much of substance. I don't trust them not to lie, but it doesn't even matter, we still need to summarize this neutrally.
How widely respected MITRE ATLAS/OWASP are is only a part of the issue. Our goal is to provide context to readers, and we need reliable sources to do this. LLM-like filler language like Security frameworks are emerging to address AI-specific development risks undermines this goal, and vague filler about how they "provide guidance" isn't helpful. To put it another way, it's not enough for you or some other editors to know what this means and why it's important, we need to use reliable, independent sources to provide this context to readers. Grayfell (talk) 20:21, 28 December 2025 (UTC)Reply
@Grayfell,
  • For the first one is the "published" IEEE SSP article in this context. I don't see a problem to cite a pre-print for a article that has been peer-reviewed. The peer-review and content is what matters, the site it is hosted on is meaningless (especially amongst the security community, where open-sourcing one's work is strongly valued and folks will host papers on their person websites for folks to read and give feedback on).
  • Wrt to the Codex paper, I'm pretty sure a good portion of the 7k citations would be relevant, and Google Scholar typically does not include random blog posts. As I said previously, the content of the paper (which is again what matters, not the site it is hosted on) is published by subject matter experts, researchers invited by OpenAI to evaluate the model. I don't see a reason for us to discount it just cause it has been published in ARXIV
  • I really don't think Anthropic AI is lying about something this simple, simply because (as a expert) there isn't anything else besides "source-code repositories" to train on and generate functional code from natural language prompts. is literally the goal of all the benchmarks (we can argue about the validity of the benchmarks, but it still stands that they are testing some subset of "generate functional code from text").
  • To your point about "knowledge", I don't get the difference you are trying to point at here. Knowledge in the context of any computer system (AI or not) is some kind of encoding of that information. I don't see what other definition Anthropic or others are pointing at, and the fact that LLMs encode this kind of information about their training dataset is a well known fact and is kinda the foundations of machine learning and deep learning at a mathematical level. I don't think Anthropic is claiming anything revolutionary here, they are saying "hey we trained on a lot of open-source code and the output reflects that data in the dataset". To that point, a quick Google search brought up which finds that LLMs do "understand" (in the LLM context, not the human sense) programming languages.
  • MITRE ATLAS and OWASP's GenAI framework (or even Snyk which currently manages much of the vulnerability disclosure framework in the Javascript ecosystem) are evaluation frameworks by organization that have a significant history in looking at vulnerability analysis and are well respected amongst the security community. (MITRE's previous ATT&CK framework is considered to foundational in cybersecurity) If you don't count them to be independent and reliable sources about security issues in LLMs, you are explicitly biasing the sourcing towards gossip rags and hype outlets like The Register and The Verge, both of which are infinitely less independent and technically inept at understanding vulnerabilities compared to the literal company that managed the CVE database (or CNAs) that is funded by the department of homeland security (note, not AI companies or venture capital) and and whoes impending lapse in funding was considered a significant event for cybersecurity. TLDR, I consider them WP:RSes in their own right, and do think that their reports need to be included in any discussion about vulnerabilities in LLM software.
Sohom (talk) 21:37, 28 December 2025 (UTC)Reply

Serious quality issues

[edit]

This article is incredibly dubious. Terminology is used haphazardly and confusingly, sources behind references do not back up the claims behind the paragraphs where they are referenced. Most sources are biased (from companies selling vibe coding-/AI assisted coding tools), and the claims made are beyond what credible sources state and what the technology is capable of. E.g. "inferring developer intent" in the section "Intelligent code completion" (should instead discuss pattern matching).

I started editing the article by correcting some issues, but it seems there's more misleading information than valuable content in the article.

The whole article reads as an advertisement. I'm on the fence between suggesting this article is deleted or that it is fused with the Vibe Coding article, which is much higher quality. TietoTeekkari (talk) 09:01, 23 April 2026 (UTC)Reply

Even more confusingly, the introduction claims LLM tools are used for *code generation*, i.e. generating machine code. This is such a wild claim I'm unsure it wasn't hallucinated. TietoTeekkari (talk) 15:25, 23 April 2026 (UTC)Reply
@TietoTeekkari : Just interested in what content you consider to be promotional of specific named products and/or services? -- JMC (talk) 23:24, 24 April 2026 (UTC)Reply
The general unsourced claims about AI-assisted software development as a technology. There are no critical sources, the article promotes the technology as a whole, usually with statements that are wildly exaggerated.
Examples:
"going beyond simple keyword matching to infer the developer's intent and picture the broader structure of the developing codebase."
Unsubstantiated claim about "matching developer intent". Completely meaningless, reads like an advert.
"Similarly, AI agents are used to perform static code analysis, identify security vulnerabilities, suggest performance improvements and ensure adherence to coding standards and best practices."
Specifically, "ensure adherence to coding standards and best practice." Unsubstantiated claim not supported by the source. Also, as a computer science researcher who has studied LLMs, this is not even possible. In fact, as far as I know, the opposite is true.
"An analysis has shown that such use of LLMs significantly enhances code completion performance across several programming languages and contexts, and the resulting capability of predicting relevant code snippets based on context and partial input boosts developer productivity substantially."
This is just plagiarized straight from the source's abstract, replacing "Our" with "An". Indicating that the editor 1) only read the abstract, 2) doesn't understand how to cite, 3) wanted a quick source to back up a claim and cherry picked an article.
Instead of a criticisms section, we have a "challenges" section, which does not mention issues of quality, which are substantial.
There is an "industry perspectives" section, only citing industry actors with vested interests in the technology.
The shoddy use of sources also indicates that this article has been written by inexperienced editors wanting to make claims fast. TietoTeekkari (talk) 17:31, 25 April 2026 (UTC)Reply
Generally, criticism sections should be avoided, but otherwise, I agree with all of these points. The use of vague filler could be a product of using LLMs to write the article, or it could just be a misguided use of LinkedIn-speak on Wikipedia. Either way, substandard writing is spread across the entire article. Grayfell (talk) 21:24, 13 May 2026 (UTC)Reply
I can accept no criticism section, although in my opinion criticism of AI-assisted programming is a notable phenomenon of its own. I'm glad someone else sees the issues with this article. I think my preferred resolution would be to salvage what can be salvaged from here and move efforts over to the vibe coding article. But I'm not against improving this article either. TietoTeekkari (talk) 13:17, 14 May 2026 (UTC)Reply
RE merging this with vibe coding - I'd be against that as vibe coding is a subset of AI software development. Whilst colloquially I agree they are often used synonymously, the sources in the vibe coding article make it clearly separate from the broader use of AI in software development. SmartSE (talk) 13:59, 14 May 2026 (UTC)Reply
I don't agree. I think there is no difference between vibe coding vs. "AI-assisted programming", and vibe coding proponents simply dislike the term "vibe coding", which they feel has connotations of "sloppy" or "lazy". That is why they are trying to create a demarkation. E.g. Simon Willison has said explicitly that this is why he wants to use a different term. His preferred term is "Agentic Engineering".
Most of the sources on this topic are also biased in this direction, from people who have vested interests in LLM programming tools. TietoTeekkari (talk) 15:30, 14 May 2026 (UTC)Reply
I'm with SmartSE on this one, seeing a clear distinction between the two approaches to software development and opposing merging of the two articles. As a practitioner, I see the essential difference being one of discipline. As the lede makes clear, in AI-assisted software development, AI is used to assist in a range of tasks of the (professional) software development life cycle. Vibe coding deliberately eschews this disciplined approach, which is not to say that it carries any emotional overtones of sloppiness or laziness. -- JMC (talk) 19:56, 14 May 2026 (UTC)Reply
I've heard this argument before, but it is (imo) an artificial distinction. I can't think of another case where a development method is called something else when discipline is involved. Some discipline is always involved. How much discipline makes something not vibe coding?
The practices are close enough that they don't understand why they would warrant two separate articles. And the vibe coding article is much, much better. TietoTeekkari (talk) 20:10, 14 May 2026 (UTC)Reply
This issue should be decided based on reliable sources, ideally independent sources, not industry blogs, and especially not first-hand experience.
From a non-practitioner's point of view, a redirect to 'vibe coding' would explain the larger topic better than this article in its current state. IMO the article is poor enough that WP:TNT is worht considering, but obviously that's pretty extreme.
Rhetorically speaking, is vibe coding a subset of this, or is this a subset of vibe coding? I could see both arguments being made, but it isn't up to us as editors to be making those arguments personally, we need to use sources. I think we all agree that there is a lot of overlap. So do reliable, independent sources treat these two concepts as distinct? What is the best way to summarize how sources talk about this concept? Would combining the articles work better? Grayfell (talk) 23:36, 14 May 2026 (UTC)Reply

I propose a root and branch rewrite of this article, in which 'AI-assisted software development' is identified as an umbrella term for any and all software development where AI is part of the process. 'Vibe coding' and 'agentic engineering' would each be recognised as subtypes of generic AI-assisted software development.

Such a rewrite would conform to the characterisation of the field by Arvind Narayanan and Sayash Kapoor in this Substack article.: 'Why AI hasn’t replaced software engineers, and won’t' (I'm not suggesting that that article be used as a RS for the rewrite itself.) The section 'Vibe coding is not agentic engineering' is particularly relevant. -- JMC (talk) 01:20, 13 June 2026 (UTC)Reply

"Agentic engineering" is not a term that sits right with me. It's a marketing term, and it's misleading.
Software engineering is an actual engineering field, "Agentic engineering" is missing the qualifier of what kind of engineering is happening. It's not a term used by anybody who isn't part of the hype.
"Agentic software engineering" I might accept, as well as "AI assisted software development/engineering", which is what this article is called. There the "engineer" title is attached yo regular software engineering, so nothing sketchy there.
But the term engineer is protected in most jurisdictions. I would feel incredibly queezy using the term "agentic engineer" (except when describing how LLM proponents use the term) in an encyclopedia. It's completely meaningless. TietoTeekkari (talk) 13:52, 13 June 2026 (UTC)Reply
I agree and further, shuffling around the names of subsections to add WP:OR about definitions is not going to fix big problem. This proposed rewrite makes it look even more like this is a WP:POVFORK of vibe coding.
Any rewrite should start with reliable, independent sources. Obviously that excludes the Substack, but it also excludes those blogs, tweets, and preprints which are cited by that Substack post. Instead of deciding what makes sense and working backwards, we should be willing to summarize critical, outside sources. This isn't a BLP, so there is no reason to spare the feelings of unnamed, hypothetical "agentic engineers' who bridle at being called vibe coders. Grayfell (talk) 19:23, 15 June 2026 (UTC)Reply

Federal Office for Information Security PDF in lead

[edit]

In this edit I moved a source from the lead to the external links section:

This source is useful, but this appears to be a selective use of the introduction of the source, citing one paragraph on p.3 while ignoring surrounding context. This subtly misrepresents the intent and purpose of the source. The source is explicitly about security, and this summary is followed by comments about poor quality control and major security vulnerabilities. Per that same page:

• AI coding assistants are no substitute for experienced developers. An unrestrained use of the tools can have severe security implications.

• A systematic risk analysis should be performed before introducing AI tools (including an assessment of the trustworthiness of providers and involved third-parties).

• Gain in productivity of development teams must be compensated by appropriate scaling measures in quality assurance teams (AppSec, DevSecOps).

• Generated source code should generally be checked and reproduced by the developers.

This is presented as the main point of the entire source. Boiling this down to saying that it helps with certain tasks is, at best, a wasted opportunity.

Grayfell (talk) 01:19, 15 May 2026 (UTC)Reply

Early full-stack example

[edit]

ChatGPT was released in November of 2022, in December and Jan 2023 I had a full stack app written entirely by ChatGPT. I had it show me how to get BLOOM to generate text of different variety, then present them to the user (ironically a way to create training material for later versions of ChatGPT). See here; https://www.youtube.com/watch?v=OwOGPw6BqtM JoeHenzi (talk) 05:12, 10 June 2026 (UTC)Reply

COI disclosure — proposing security content

[edit]

I am disclosing a conflict of interest before proposing changes to this article.

I have a reputational interest in this paper being cited. I will not edit the article directly and am proposing changes here for community review. HEHwIeVfk (talk) 20:52, 2 July 2026 (UTC)Reply

Edit request: Security content

[edit]

A 2024 empirical study analyzed 7,703 source code files from public GitHub repositories that were explicitly attributed to AI code generation tools via developer comments in the source code, using CodeQL static analysis to identify security vulnerabilities.[1] Files without explicit attribution comments were excluded, meaning the dataset captures cases where developers openly acknowledged AI assistance rather than the full extent of AI tool usage. The study found that 87.9% of analyzed files contained no detectable Common Weakness Enumeration (CWE)-mapped vulnerabilities. Vulnerability rates varied significantly by programming language. Python files showed rates of approximately 16–18%, JavaScript 8–9%, and TypeScript 2–7%. This language-dependent pattern remained consistent across all four AI tools examined (ChatGPT, GitHub Copilot, Amazon CodeWhisperer, and Tabnine), suggesting that programming language characteristics influence security outcomes more strongly than the choice of AI tool.[1] The five most critical vulnerability types identified, ranked by average CVSS v3 severity score, were SQL injection (CWE-89, avg. 8.76), OS command injection (CWE-78, avg. 8.68), code injection (CWE-94, avg. 8.68), use of hard-coded passwords (CWE-259, avg. 8.64), and use of hard-coded credentials (CWE-798, avg. 8.57). Three of these five types (SQL injection, OS command injection, and code injection) appear in MITRE's 2025 Top 25 Most Dangerous Software Weaknesses.[1][2]

[1] Schreiber, Maximilian; Tippe, Pascal (2025). "Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis of Public GitHub Repositories". Information and Communications Security: 27th International Conference, ICICS 2025. Springer. pp. 153–172. doi:10.1007/978-981-95-3537-8_9. A freely available preprint is available at arXiv:2510.26103.

[2] MITRE (2025). "CWE Top 25 Most Dangerous Software Weaknesses". https://cwe.mitre.org/top25/. Retrieved 2026-07-02.


HEHwIeVfk (talk) 21:12, 2 July 2026 (UTC)Reply

Reply 3-JUL-2026

[edit]

🔼  Specification requested  

  • It is not known what changes are requested to be made. Please state your desired changes in the form of "Change X to Y using Z". (See WP:CHANGEXY.)

Kindly open a new edit request at your earliest convenience when ready to proceed.
Regards,  Spintendo  03:20, 4 July 2026 (UTC)Reply

I don't think I've heard of "27th International Conference, ICICS 2025" (and I consider myself fairly well versed in computer security in this realm) Sohom (talk) 07:47, 4 July 2026 (UTC)Reply
Seems to be the 2025 International Conference on Information and Communications Security, October 29-31 2025, held in Nanjing, China. From the Conference Program (available online): "ICICS was initiated in 1997, this year marks the 27th anniversary of ICICS conference. The goal and feature of this conference is to bring together researchers and practioners [sic] from both academia and industry to discuss and exchange their experiences, lessons learned, and insights related to information and communications security. The program this year was comprised of 3 keynote lectures and 18 oral sessions." -- JMC (talk) 20:15, 4 July 2026 (UTC)Reply