What counts as plagiarism when nothing is copied word for word?

Most people picture plagiarism as a side-by-side comparison: two paragraphs, the same sentences, perhaps a clumsily swapped synonym here and there. It is a tidy mental model, and it is the one that copy-detection software was built to enforce. It is also, unfortunately, wrong – or at least so incomplete that it misleads almost everyone who relies on it.

The academic and publishing worlds have understood for decades that plagiarism is fundamentally a question of intellectual debt rather than textual overlap. The Office of Research Integrity in the United States defines it as “the appropriation of another person’s ideas, processes, results, or words without giving appropriate credit”, a formulation that puts ideas before words quite deliberately. The Modern Language Association is blunter still:

“Presenting another person’s ideas, information, expressions, or entire work as one’s own is […] plagiarism.”

Notice what is missing from both definitions. Neither requires that any sentence be copied. Neither sets a percentage threshold. Neither mentions string matching at all. Yet the tools and habits that have grown up around plagiarism enforcement, from undergraduate similarity reports to journalistic accountability, have quietly inverted the priority – treating textual identity as the thing being policed and intellectual debt as a fuzzy afterthought.

This piece is an attempt to put the priorities back in the right order. If you only remember one thing from what follows, let it be this: plagiarism is the failure to acknowledge a source debt, and a source debt can be incurred without ever touching the source’s prose.

The five things you can actually steal

It is useful to break “intellectual debt” into its constituent parts, because each generates its own variety of plagiarism and each is invisible to a naive textual comparison.

1. Ideas

The most fundamental form of unacknowledged borrowing is the idea itself – a hypothesis, an interpretation, a theoretical framework, a novel claim about how the world works. If a researcher reads a paper proposing that, say, urban tree canopy density predicts childhood asthma rates, and then writes their own paper advancing that thesis without citing the original, they have plagiarised even if every word of their manuscript is freshly composed. The debt is to the conceptual move, not to the language used to make it.

This is why the Council of Science Editors and most journal style guides explicitly distinguish “plagiarism of ideas” as a category in its own right. It is also why peer reviewers are bound by confidentiality: a reviewer who reads an unpublished manuscript and then incorporates its central insight into their own work – even rewritten from scratch – is committing perhaps the most serious form of academic theft, sometimes called “idea plagiarism” or, more pointedly, scooping.

2. Argument structure

Beneath the surface of any non-trivial piece of writing sits a scaffolding: the order in which claims are introduced, the way evidence is marshalled, the particular sequence of objections raised and dispatched. This scaffolding is itself an intellectual artefact. A philosopher who, say, refutes a position by first conceding its strongest version, then identifying a hidden premise, then producing a counterexample that targets that premise specifically, has built something. Reproducing that architecture in your own words, against the same opponent, without citation, is a form of plagiarism that no string-matching algorithm will catch.

The legal academic Richard Posner, writing in The Little Book of Plagiarism, made this point with characteristic sharpness: what readers value in scholarly work is often “the sequence and structure of the argument” rather than the prose itself, and so it is precisely the sequence and structure whose theft most damages the original author.

3. Sequence and selection

Closely related, but worth distinguishing, is the question of what gets included and in what order. A historian who selects a particular set of six primary sources, arranges them chronologically against a thematic backdrop, and draws out a specific tension between them has done original intellectual labour in the selection and sequencing itself. A second historian who reproduces those choices – same six sources, same order, same tension highlighted – has borrowed something real, regardless of whether their prose is entirely their own.

This is the principle behind copyright’s “selection and arrangement” doctrine, which protects compilations even when the underlying elements are not themselves protectable. Plagiarism norms run parallel to, but are stricter than, copyright: you can plagiarise public-domain Shakespeare quotations by selecting and arranging them in a way that mirrors someone else’s anthology.

4. Method and process

In technical fields, the procedure by which a result was obtained is often more valuable than the result itself. A novel experimental protocol, a particular sequence of analytical steps, an unusual computational pipeline — these are intellectual contributions. Re-implementing a method from scratch, in different code or with different reagents, does not erase the debt to the person who designed it. The Committee on Publication Ethics treats unattributed methodological borrowing as a distinct subcategory of research misconduct precisely because the absence of textual overlap so often misleads detection systems.

5. Attribution chains

The last category is the subtlest, and perhaps the most common in the wild. Suppose author A makes a claim and cites primary source B. Author C reads A, lifts both the claim and the citation to B, and presents the result as their own reading of B – without ever opening B and without citing A. This is sometimes called “secondary citation plagiarism” or, more colloquially, citation laundering. The text may be entirely original. The references may all check out. But the path through the literature has been stolen, and with it the implicit claim that the author has done the reading.

Why “patchwriting” sits in the middle

Between flagrant copy-paste and pure idea theft lies a vast grey zone that the composition scholar Rebecca Moore Howard usefully named patchwriting: the practice of rewriting a source sentence by sentence, swapping vocabulary and rearranging clauses, while preserving the underlying structure and substance. Patchwriting is interesting because it sits exactly on the fault line between textual and structural plagiarism. A similarity report may show only modest overlap – perhaps 8% or 12%, well within most institutional thresholds – while the document is, in every meaningful sense, a derivative of its source.

Howard’s research, conducted across decades of student writing, suggests that patchwriting is often a developmental stage rather than a deliberate fraud: writers who do not yet feel authoritative in a discourse community cling to the syntax of their sources as a kind of scaffolding. This is a useful framing for educators. It is not, however, a defence. The finished document still represents a source debt that has not been declared, and treating it as acceptable simply because the prose has been sufficiently disguised teaches exactly the wrong lesson.

The “common knowledge” exception, and where it actually stops

A common defence against accusations of unacknowledged borrowing is the appeal to common knowledge: nobody, after all, cites a source for the claim that Paris is the capital of France. This is a real and important exception, but it is much narrower than people assume.

Common knowledge, in scholarly practice, means facts that an educated reader in the relevant field would already know and could verify in any of several standard reference works. The boundary moves with the field. That the Battle of Hastings occurred in 1066 is common knowledge in any historical context; that William’s army included a specific proportion of Breton cavalry is not, even though both might appear in the same encyclopaedia entry. When in doubt, the operative principle is the one articulated in the Publication Manual of the American Psychological Association: cite. Over-citation is a stylistic blemish; under-citation is an ethical failure.

Crucially, the common-knowledge exception applies to facts, not to interpretations. That a recession occurred in 2008 is common knowledge. That it was caused by a particular interaction between regulatory capture and securitisation incentives is somebody’s argument, and that somebody deserves the citation.

The shape of a source debt

It may help to think about the question backwards. Rather than asking “did I copy this?”, a more useful self-audit is: what would this piece of work look like if the source did not exist? If the answer is “broadly the same, with different examples”, you probably owe an acknowledgement but not necessarily a quotation. If the answer is “I would not have framed the problem this way”, or “I would not have selected these cases”, or “I would not have known to reject this counter-argument”, then the source has done real work for you, and the debt is substantive regardless of the surface text.

This framing – borrowed loosely from the philosopher of science Helen Longino’s account of how knowledge accumulates through community – captures something that percentage-based detection cannot. Plagiarism is not a property of a document. It is a property of the relationship between a document and the intellectual landscape from which it emerged.

What this means for detection

The implications for those of us thinking about how plagiarism is identified, by humans and by software, are significant.

First, similarity scores are a floor, not a ceiling. A document with 2% textual overlap may be entirely honest or deeply derivative; the number cannot tell you which. The serious work of plagiarism assessment begins where the similarity report ends, in the comparison of structure, sequence, and substance.

Second, the rise of large language models has made textual overlap an even weaker signal than it used to be. A model can be prompted to reproduce the argument of a source paper in completely fresh prose, and no n-gram matcher will flag it. This is not a flaw in the models so much as a vindication of what plagiarism scholars have been saying for half a century: the problem was never really the words.

Third, and most practically, the cultivation of citation habits – generous, specific, slightly paranoid citation habits – does more to prevent plagiarism than any detection tool. A writer who cites the source of every non-obvious idea, every borrowed structural choice, every methodological inheritance, simply cannot plagiarise inadvertently, regardless of how their prose happens to resemble or diverge from their sources.

A working definition

If a single working definition is wanted, this one will serve: 

Plagiarism is the presentation, as one’s own, of any element of intellectual labour textual, conceptual, structural, methodological, or bibliographic – that originated with someone else and that the audience would, if they knew, expect to be credited.

The audience-expectation clause is doing real work in that sentence. It is what distinguishes a ghost-written corporate report (where the audience has no expectation of personal authorship) from a ghost-written PhD thesis (where they emphatically do). It is what makes the same sentence plagiarism in a journal article and unremarkable in a press release. And it is what should anchor any thoughtful approach to detection, attribution, and the broader culture of credit in which honest writing has always operated.

Copy-paste plagiarism is the easy case – the one we built our tools around because we could. Everything more interesting is happening elsewhere.

References and further reading:

  • Council of Science Editors, 2018. CSE’s white paper on promoting integrity in scientific journal publications. 2018 update. Wheat Ridge: Council of Science Editors.
  • Howard, R.M., 1999. Standing in the shadow of giants: plagiarists, authors, collaborators. Stamford: Ablex.
  • Longino, H.E., 1990. Science as social knowledge: values and objectivity in scientific inquiry. Princeton: Princeton University Press.
  • Modern Language Association, 2024. Plagiarism and academic dishonesty. [online] Available at: https://style.mla.org/plagiarism-and-academic-dishonesty/(opens in new tab).
  • Office of Research Integrity, n.d. Definition of research misconduct. [online] U.S. Department of Health and Human Services. Available at: https://ori.hhs.gov/definition-misconduct(opens in new tab).
  • Posner, R.A., 2007. The little book of plagiarism. New York: Pantheon.
  • Committee on Publication Ethics, 2019. Core practices. [online] Available at: https://publicationethics.org/core-practices.