# Garbage Collection Is a Professional Skill

## Maintenance after code became cheap

amaaov · 8 September 2026

In 1960, John McCarthy described a Lisp system running on the IBM 704 with a peculiar storage problem. Programs kept requesting fresh storage even though some structures already in memory had become inaccessible. His procedure began with the registers that remained accessible, followed their chains, and returned everything else to a list of free storage; elegance, provenance, effort, and meaning had no place in the decision, which concerned only whether the running program could still reach an object. [Recursive Functions of Symbolic Expressions and Their Computation by Machine, Part I](https://doi.org/10.1145/367177.367199)

The procedure became known as garbage collection, and it can be exact inside memory because the collector observes a defined graph at a particular instant. Public software offers no equivalent graph: a repository ignored today may become evidence tomorrow, fifty lines of script may contain the missing method behind a scientific result, and an unmodified library may remain installed in thousands of systems. Meanwhile, a project receiving commits every hour may have no users at all.

Garbage is therefore a dangerous name when it turns dislike into a property of an object, but a useful one when it directs attention to relations: which artifacts occupy shared capacity, which ones invite reliance, where they transfer work, and who has accepted the task of removing or classifying them.

Generative systems have reduced the cost of producing code in many settings, while leaving the subsequent work in place. A program can now appear before anyone has decided who understands it, who will answer questions about it, how its claims will be checked, or what will happen when support stops; production has become cheaper, but answerability still costs time.

## The obligation after publication

“AI slop” commonly names digital material made with generative artificial intelligence and judged low in quality, effort, or meaning. The term is too unstable for policy because origin says little about the artifact, while quality, effort, and meaning remain unspecified: a careful person can publish a poor patch, a generator can contribute to a durable one, and even the author may be unable to account reliably for every influence and every minute of work.

[LLMs as Probabilistic Medium](https://amkisko.github.io/posts/20251111160000_llms_probabilistic_medium.html) described a generated draft as plausible before it becomes true and located the useful operator work in its passage toward a repository or customer. That earlier account concerned the production of a candidate; the obligation considered here begins when someone asks another person or a shared system to carry it.

Nobody inspecting a repository can see care directly, and time spent offers a poor substitute because long labour can produce a brittle system while a brief correction can be exactly right. Once a file enters a package registry, it also enters conventions shared by publishers, installers, maintainers, attackers, archivists, and people arriving through search. Even silence communicates there, since readers may infer that a project is supported or suitable for use when its publisher has established neither.

Inspection reaches evidence that intention cannot supply. It can establish whether the artifact does what it claims, states its readiness, and allows another person to reproduce the result, identify dependencies, report a problem, or withdraw safely; it can also expose disproportionate demands on storage, computation, bandwidth, review time, and attention, along with the share of verification and repair accepted by the publisher.

On that evidence, slop can be defined without reference to its tool of composition: it is an artifact presented for reliance while its avoidable costs of understanding, verification, operation, or repair are transferred to other people and its condition is misrepresented. The definition covers a generated patch, but it also covers a handwritten package with invented benchmarks, an unanswered security queue, and a README that promises production readiness.

The [Manifesto for Collaborative Concurrent Extreme Software and Product Development for humans](20260608120000_manifesto_collaborative_concurrent_extreme.html) once proposed adopting “slop as raw ore.” The phrase belongs to an admitted intermediate stage, where somebody has agreed to refine the material before placing it in a structure that other people must trust.

“Avoidable” and “reasonable” still require judgment because purpose and capacity differ, yet they direct that judgment toward evidence that people can discuss. A platform can measure storage, automated jobs, network traffic, abuse reports, and staff time, while maintainers can ask contributors and users about review burden, trust, support expectations, and reasons for leaving. Such observations cannot disclose an author’s feeling, but together they describe the consequences shared with others.

## Degrees of reliance

Publication is often described as the single act of putting something on the internet, although a notebook, a research deposit, a package, a contribution, and a service each establish a different relation with the people who encounter them.

A public notebook may declare itself an experiment preserved for inspection with no support promised, in which case truthful context, a licence, provenance, and basic secret and storage hygiene carry most of its obligation. A research artifact may require more exact preservation: the [Software Citation Principles](https://doi.org/10.7717/peerj-cs.86) call for persistent identification and version specificity because software can be a research product, while [software deposit guidance](https://zenodo.org/records/1327310/files/SoftwareDepositGuidance.pdf?download=1) extends that reasoning to small shell and R scripts. Size and glamour do not determine whether code belongs in a scholarly record, and a finite artifact may be complete. [Existence](20260827165000_existence_en.html) approaches the same boundary from the contributor’s side, where a deposited file can remain available without making its contributor responsible for everything subsequently derived from it; the public description must keep that limit legible.

A reusable package asks more because, once a release enters an installer, other systems can depend on its name and version without their operators returning to its homepage. Compatibility, supported versions, security contact, ownership, and a withdrawal path become part of its interface. A pull request enters somebody else’s queue and consumes limited review time through reproduction, explanation, revision, and possible withdrawal, while a hosted service adds operational capacity, incident response, privacy, and continuity. The reliance induced and the work invited from other people determine the obligation.

This distinction matters in the dispute around Codeberg’s essay [Protecting our FLOSS commons from LLMs](https://blog.codeberg.org/protecting-our-floss-commons-from-llms.html), which describes autonomous or resource-heavy generated projects as a threat to a volunteer-run commons while saying that small experimental side projects using few resources will probably be tolerated. The surrounding documentation identifies concrete limits: default storage allows 750 MiB for Git repositories and another 1.5 GiB for packages, releases, attachments, and large-file storage; continuous integration consumes substantial resources; and large binaries remain expensive in repository history. [Storage limits](https://blog.codeberg.org/new-storage-limits-on-codeberg-what-you-need-to-know.html), [continuous integration guidance](https://docs.codeberg.org/ci/), [large-file guidance](https://docs.codeberg.org/git/using-lfs/)

These limits justify governance of shared capacity, although the cited public materials provide no distribution that separates ordinary source repositories from automated mirrors, binary stores, continuous-integration workloads, popular downloads, abuse handling, and exception review. Operators may still need discretion, but a public argument becomes more defensible when it names the behaviour being restricted and publishes whatever evidence can safely be shared.

A forge described as a commons remains a bounded service with admission rules, finite maintenance, and people empowered to protect it. Its rules should describe conduct that participants can understand, and their enforcement should leave room for appeal, correction, and legitimate experiments.

## The older maintenance problem

The present dispute returns to questions that were already explicit in 2022. On 3 February of that year, the US National Institute of Standards and Technology published version 1.1 of its [Secure Software Development Framework](https://doi.org/10.6028/NIST.SP.800-218), covering preparation, release integrity, vulnerability response, and correction of underlying causes throughout software development and maintenance. On 3 May, NIST published [open-source software controls](https://www.nist.gov/itl/executive-order-14028-improving-nations-cybersecurity/software-supply-chain-security-guidance-22) that treated provenance, integrity, support, and maintenance as important qualities that were often difficult to discover. ChatGPT’s public research preview arrived on 30 November. [Introducing ChatGPT](https://openai.com/index/chatgpt/)

This chronology matters because conversational generation widened access to software production without originating the maintenance deficit, the insecure dependency, the unreadable build, or the absent owner. Those problems had accumulated throughout software development and maintenance, whereas current arguments often compress them into an attempt to classify the first draft.

A 2022 study of 265,325 pull requests across ten established open-source projects provides one empirical view of that earlier problem. The researchers identified 4,450 proposals abandoned by their contributors and associated abandonment with contributor, review, workload, and project factors; although the study neither measures generative code nor represents every project, it shows how accepting proposed work has long depended on human participation after submission. [On Wasted Contributions](https://doi.org/10.1145/3530785)

[When Reading Costs More Than Writing](https://amkisko.github.io/posts/20260313103000_when_reading_costs_more_than_writing.html) records the same inversion at the scale of one review, where verifying a generated commit message cost more than replacing it by hand. A small qualitative study published in 2024 moves outward from that scene: interviews with ten maintainers from nine well-adopted projects describe maintenance labour as depletable and project continuity as dependent on human infrastructure. The sample cannot establish prevalence across open source, but it does show why a repository cannot be understood as code plus an issue tracker; the working arrangement also includes attention, relationships, conflict handling, succession, and rest. [A 2024 study of maintenance labour and human infrastructure](https://arxiv.org/abs/2408.06723)

When a new tool multiplies proposals faster than communities can inspect them, judgment becomes the constrained resource. Limits, queues, deposits, automated checks, contributor duties, and refusal may all protect that resource, whereas a ban defined only by composition method can overlook an expensive handwritten proposal.

## Proportion before publication

The instruction to optimize as much as possible before publication begins from a sound demand for cleanup, but optimization has no universal maximum. Runtime, memory, energy, code size, accessibility, security, maintainability, reviewer time, and time to feedback can conflict, and the [ISO software product quality model](https://www.iso.org/standard/78176.html) contains nine broad characteristics that the word *efficient* cannot settle by itself. The [Software Carbon Intensity specification](https://sci.greensoftware.foundation/) likewise needs operational energy, grid intensity, embodied hardware emissions, and a functional unit before it can express one kind of efficiency, while a faster request says little about total resource use if the number of requests changes.

Donald Knuth’s famous optimization passage is usually compressed until its practical instruction disappears. In context, he criticized spending time on noncritical parts of programs and urged measurement to identify the critical parts where improvement is worthwhile. [Structured Programming with go to Statements](https://dl.acm.org/doi/10.1145/356635.356640) Open development also has a countertradition of releasing early because users reveal requirements that private polishing cannot. A 2013 analysis of SourceForge projects found a curved relationship between release frequency and later download share, with increases helping only to a point before they could backfire; the historical sample and its download metric establish no universal schedule, but they explain why preparation and release frequency both need stopping rules. [Release Early, Release Often?](https://openreview.net/forum?id=cBWCD2kJP6)

A proportionate instruction follows from these limits: before asking a public system or another person to carry the work, remove the avoidable costs that can be identified, measured, and afforded, then publish at the earliest level of readiness that can be described honestly. For a notebook, this may require removing secrets and generated binaries before adding a brief account of purpose and status; a contribution may require reproducing the defect, running focused checks, explaining the trade-off, and remaining available for review; a service may require load tests, resource budgets, rollback, monitoring, and an operator who can respond. These stated and tested limits provide an attainable standard when perfection cannot.

Re-doing an application can be part of that preparation when the second implementation is used to recover knowledge. A team rebuilding an in-house tool must follow the work its old labels compressed: who may act, which inputs arrive, which exceptions matter, how a failed operation is corrected, and what another system expects next. Differences among the old behaviour, the written instruction, and a colleague's actual route reveal facts to preserve or choices to change. Tests can hold the behaviour that must survive, while user documentation and runbooks gain examples and recovery paths encountered in use.

In 2024, Plain reported a small version of this practice in its own support platform. Its team built stand-alone applications for common customer uses against the company's API; the work exposed gaps in the API, documentation, and developer experience while leaving code examples that could be shared. This is productive duplication, where the extra program serves as an instrument for making the product's account more complete. The evidence remains narrow: Plain also reported customers finding scale problems that its own usage did not reveal. [Dogfooding at Plain](https://www.plain.com/blog/dogfooding-at-plain)

Rebuilding can also destroy the evidence accumulated in the first implementation. Joel Spolsky's warning about full rewrites observes that apparently awkward code may preserve bug fixes discovered only after real-world use, so discarding the implementation can discard hard-won knowledge with it. A second application earns value by recording those discoveries, comparing behaviour under actual work, and retiring the duplicate deliberately. That passage through use and transfer keeps the replacement from becoming another undocumented claim in the pile. [Things You Should Never Do, Part I](https://www.joelonsoftware.com/2000/04/06/things-you-should-never-do-part-i/)

Cleanup inside deployed software also has a different meaning from the classification of a public artifact. An exploratory study of twenty-three open-source Java desktop applications found that methods classified as unused generally persisted for long periods, were rarely restored to use, and were often already unused when introduced. The selected projects, analysis tools, and treatment of reflection limit the conclusion, yet the distinction remains useful: code that takes no part in a deployed implementation can obstruct understanding, while an inactive artifact placed honestly in an archive may continue to serve as a record. [Study of unused methods in open-source Java desktop applications](https://link.springer.com/article/10.1007/s10664-023-10303-0)

## Deprecation and preservation

The publisher of software can maintain it, transfer it, deprecate it, or archive it, and each choice can be expressed through ordinary technical mechanisms. [GitHub](https://docs.github.com/en/repositories/archiving-a-github-repository/archiving-repositories) can make an archived repository read-only and recommends closing open work and updating its description or README first. [npm](https://docs.npmjs.com/deprecating-and-undeprecating-packages-or-package-versions/) can attach a deprecation message to an installed package and describes transfer when a depended-upon project needs a new owner. The Apache Software Foundation moves projects whose development has stopped into its [Attic](https://attic.apache.org/), preserving their material while stating their condition.

Python package indexes now have a more precise vocabulary in the accepted [Project Status Markers specification](https://packaging.python.org/en/latest/specifications/project-status-markers/), which defines active, archived, deprecated, and quarantined states. A project without an explicit marker is treated as active; archived projects reject new uploads while retaining existing distributions, deprecated projects can direct installers elsewhere, and quarantine is reserved for unsafe material. The status is machine-readable because a banner on a repository page may never reach the person typing an installation command. In a dependency system, an omitted marker leaves the project in the active state, so any withdrawal should be announced where reliance occurs.

Preservation is another responsible destination. [Software Heritage](https://www.softwareheritage.org/mission/) seeks to collect and preserve publicly available source code as scientific, technical, and cultural knowledge, thereby cautioning against the use of present activity as the only measure of value. A memory collector can reclaim an inaccessible object because it knows the graph and the instant; historical value cannot be decided from such a fixed set of starting references.

Mary Douglas’s work is often condensed into the idea of dirt as matter out of place, although later scholarship notes that the slogan’s earlier attribution is uncertain. That uncertainty suits an argument concerned with maintenance, while the enduring point remains relational: disorder appears within a scheme of classification. A prototype may be properly placed in an archive and dangerously misplaced when advertised as a supported security library. [Placing Matter Out of Place](https://www.tandfonline.com/doi/full/10.1080/13264826.2013.785579)

Steven Jackson’s “broken-world thinking” extends the argument by locating technological creativity in repair and maintenance. [Rethinking Repair](https://doi.org/10.7551/mitpress/9780262525374.003.0011) A repairer encounters dependencies, handoffs, informal knowledge, unequal burdens, and cases in which restoration requires transformation; this work belongs to the making of shared technical arrangements because every new object arrives among systems already subject to failure and revision.

## The person who answers

The counsel to “put love into what you do” belongs to the person doing the work, where feeling may guide attention and patience. A repository offers no proof of that feeling, although readers can encounter its practical consequences in a bounded claim, a reproducible example, an answered question, a release that can be rolled back, a clear warning, a patient handoff, or an archive notice.

Care for technical work begins with the people asked to use, review, or maintain it. Depending on the circumstances, that care may mean supporting a program, declining a feature so maintainers can rest, funding repetitive triage, reducing continuous-integration demand on a volunteer forge, writing a migration guide before support stops, or replacing an ambiguous project status with an explicit one.

Communication gives the obligation an address. An author’s account can explain intention, while the artifact’s behaviour shows what users encounter, and meaning develops through the relation among people, conventions, and consequences. [Duty Before the Answer](20260816205000_duty_before_answer_en.html) places this obligation at the moment a person chooses whether to stand behind a diff; maintenance carries it forward through subsequent releases, questions, transfers, deprecation, and archival.

Garbage collection, in this broader and necessarily imperfect sense, is a professional skill for anyone who asks others to depend on published work. It requires removing unused branches, closing obsolete queues, measuring expensive operations, marking experiments, transferring ownership, and preserving records where they can be found, all in service of reducing avoidable burdens and clarifying which interfaces and commitments remain supported.

McCarthy’s collector could stop the Lisp program, trace the structures that remained accessible, and recover the rest. Public software permits neither a comparable pause nor a permanent map of social reachability, which marks the boundary of the analogy: a maintainer must ask who still depends on an artifact, which promises remain in force, what history deserves preservation, and who bears the cost of keeping it available.

Whatever tools assisted its production, public code makes a claim about future use. Professionalism begins when the person making that claim remains answerable for its consequences and communicates clearly when that answerability changes form or comes to an end.
