ACM Queue

ACM Queue

Share

Queue is the ACM's magazine for practicing software practitioners. Queue does not focus on either industry news or the latest "solutions."

Queue focuses on the technical problems and challenges that loom ahead, helping readers to sharpen their own thinking and pursue innovative solutions. Rather, Queue takes a critical look at current and emerging technologies, highlighting problems that are likely to arise and posing questions that software engineers should be thinking about.

A Valid Tool Call Can Still Be an Invalid Effect | Queue 09/28/2026

A Valid Tool Call Can Still Be an Invalid Effect

Why consequential agent actions need a semantic read set

Agent tool protocols make functions discoverable through names, descriptions, and argument schemas. That is enough to form a request but not enough for a consequential system to decide whether the resulting effect is still acceptable. Traditional software usually fixes a transaction pattern in code: which state is read, which authority is consulted, which conditions are checked, and which target is updated. An agent can assemble that pattern at runtime and emit only the final write, allowing the material read set to disappear into model context, tool transcripts, or logs.

This article calls this interface failure read-set erasure, defines the missing information as the effect’s semantic read set, and proposes a dependency-carrying effect: a canonical effect, target-required dependencies, bounded authority, replay semantics, and an explicit commit topology. The agent may propose dependencies, but the protected effect type and target must define the checks that cannot be omitted. A detached pre-dispatch verifier is not atomic; systems must state whether they provide target-native conditional commit, reservation-bound ex*****on, or saga-bound dispatch. The mechanisms are familiar: conditional writes, optimistic validation, reservations, idempotency, and sagas. The contribution is an interface discipline that preserves the justification for an agent-generated effect until the system of record can accept, hold, or refuse it.

A Valid Tool Call Can Still Be an Invalid Effect | Queue Why consequential agent actions need a semantic read set Agent tool protocols make functions discoverable through names, descriptions, and argument schemas. That is enough to form a request but not enough for a consequential system to decide whether the resulting effect is still acceptable. Traditio...

CAFE(S): Your Agent Is Only As Good As Its Context | Queue 09/24/2026

CAFE(S): Your Agent Is Only As Good As Its Context

Five durable properties for evaluating context quality

As organizations delegate increasingly complex software tasks to AI agents, engineering leaders often attribute system failures to model capabilities or harness orchestration. Even state-of-the-art models, however, degrade rapidly when operating on unclear, incomplete, or stale inputs. We present CAFE(S), a diagnostic-quality framework that evaluates assembled context across five core dimensions: clarity, actionability, fidelity, efficiency, and security. CAFE(S) provides a shared vocabulary for identifying recurring context failures—from ambiguous requirements and impossible tasks to excessive noise and unsafe inputs—and for reasoning about how those failures affect human-agent work. We argue that context quality is a first-class engineering concern that can be deliberately designed, evaluated, and maintained. Organizations that systematically strengthen the information environments in which agents operate will be far better positioned to realize their full potential.

CAFE(S): Your Agent Is Only As Good As Its Context | Queue Five durable properties for evaluating context quality As organizations delegate increasingly complex software tasks to AI agents, engineering leaders often attribute system failures to model capabilities or harness orchestration. Even state-of-the-art models, however, degrade rapidly when operating...

Finding the Real Bottleneck in Large-Scale ML Training | Queue 09/21/2026

Finding the Real Bottleneck in Large-Scale ML Training

Real efficiency requires fixing the underlying systems

Planning meetings across AI infrastructure teams are routinely dominated by accelerator provisioning and physical supply shortages. An equally critical and often overlooked constraint is operational efficiency. In production training clusters, actual model ex*****on frequently consumes only half of an accelerator’s theoretical capacity. This 50 percent utilization drag is driven by compounding system overheads. When cluster utilization drops, runtimes increase and effective ex*****on costs double even when model architectures and hardware remain identical. This article deconstructs the primary sources of compute loss in large-scale distributed ML systems, evaluating real-world scaling degradation across MLPerf benchmark ceilings and multitenant production traces. It demonstrates why rigorous workload placement discipline, bare-metal AI topologies, and establishing Model FLOPs Utilization (MFU) as a primary operational SLA must be made a priority alongside hardware procurement to control total cost of ownership (TCO).

Finding the Real Bottleneck in Large-Scale ML Training | Queue Real efficiency requires fixing the underlying systemsPlanning meetings across AI infrastructure teams are routinely dominated by accelerator provisioning and physical supply shortages. An equally critical and often overlooked constraint is operational efficiency. In production training clusters, ac...

queue.acm.org 09/17/2026

Signs of a Brewing Crisis

Crises have the destabilizing power to shift incentives and invert power structures.

Crisis is a label often applied to the conditions of an organization’s complex systems by practitioners and by managers to coerce prioritization. Infrequently, it is used to describe a situation nearest its definition: a crucial turning point of change, which requires adaptive resilience for procession.

There are technical and organizational practices that are effective, or available, only during a real crisis. With a combination of existing research, lived experience, and real-world examples, this article outlines a set of five criteria useful for evaluating the presence or imminence of a crisis: fundamental surprise, failure of sensemaking, disruption of core processes, high visibility, and a rigid deadline.

Ability to identify and evaluate these circumstances can allow operators and engineers to anticipate opportunities when adaptive resilience will be required—when a different set of technical features and organizational behaviors will be most successful in the management and operations of complex systems.

queue.acm.org

What’s in Your Software Repair Kit? | Queue 09/14/2026

What’s in Your Software Repair Kit?
The things you wish you had before a crisis

You can’t prevent or predict the next systems meltdown. But you can build resilience into your systems and organization before the next crisis strikes. Drawing on decades of site reliability engineering (SRE) experience and dozens of engagements with organizations in crisis, this article suggests technical and organizational patterns that will give your systems better odds of surviving the unexpected.

What’s in Your Software Repair Kit? | Queue The things you wish you had before a crisisYou can’t prevent or predict the next systems meltdown. But you can build resilience into your systems and organization before the next crisis strikes. Drawing on decades of site reliability engineering (SRE) experience and dozens of engagements with ...

queue.acm.org 09/09/2026

Data Architectures in the Age of AI

Every computing era creates a new software stack

Every major computing shift has reshaped the software stack, but one assumption held across all of them: Humans were the primary consumers of software.

AI agents break this assumption. An autonomous agent determines at runtime what data it needs, interprets it, and takes action, with no human in the loop to catch misinterpretations.

That change lands hardest on the data layer. Today’s platforms are adding AI capabilities to existing warehouse-, lake-, and application-centric architectures. But bolting AI features onto existing stacks isn’t enough. The data layer must be rethought from first principles to meet the requirements of autonomous agents.

This article identifies and discusses three imperative changes: Business semantics must be machine readable and continuously synchronized, federation must become the default for operational workflows, and governance must shift from controlling access to governing autonomous behavior. These changes are essential to building effective AI agents.

queue.acm.org

queue.acm.org 09/09/2026

How CHERIoT Provides Strong and Able Isolation Without an MMU

Co-design of hardware and software offers stronger security and ease of use

CHERIoT is optimized for small devices but showcases what is possible at any scale if you start by assuming CHERI.

queue.acm.org

When You Wish Upon a Prompt | Queue 09/04/2026

Kode Vicious
When You Wish Upon a Prompt
Relying on an LLM doesn’t teach how to write code

With the current LLMs, anyone can generate a working program, but a program is not a system. Today’s students must be taught not how to spot an errant asterisk, but how all the components of a system work together. This means more mentoring, debugging large working systems, and asking hard questions about how a system looks when it is properly put together.

When You Wish Upon a Prompt | Queue Relying on an LLM doesn’t teach how to write code With the current LLMs, anyone can generate a working program, but a program is not a system. Today’s students must be taught not how to spot an errant asterisk, but how all the components of a system work together. This means more mentoring, debu...

queue.acm.org 08/25/2026

Arithmetic Without Numbers

What happens inside an LLM when it tries to calculate with nothing but matrices

How does a large language model do arithmetic when all it has are matrices—no fingers, no scratch paper, no columns of digits? This article looks inside a frozen Llama and asks a sharper question than whether the model can call a calculator: with the prompt text removed, can its own activations reveal the operation and operands? In these experiments, they can. Under a strict no-parser rule, activation-derived readouts supplied a calculator’s arguments for several arithmetic tasks. The same activations reveal number-related directions and rotations, a geometry closer to a clock or spiral than to written columns. Interventions test whether selected internal states affect behavior, while provenance audits trace the calculator arguments back to model activations. Together, the results show one way arithmetic structure can take shape inside a language model.

queue.acm.org

Too Many DevEx Metrics, Too Little Guidance - ACM Queue 08/18/2026

Too Many DevEx Metrics, Too Little Guidance

DevEx Metrics Compass: making sense of developer experience measurement

As AI-augmented development becomes standard practice, engineering leaders face mounting pressure to demonstrate its impact. Yet most organizations are measuring AI adoption and output while the effect on developer experience (DevEx) remains largely unknown. The right metrics can close that gap, surfacing the everyday friction developers face and monitoring the real impact of initiatives such as AI rollouts. But with more than 120 metrics in circulation across frameworks, companies, and scientific literature, identifying the right, meaningful, actionable ones for a specific context has become a challenge in itself.

This article introduces DevEx Metrics Compass, a public, open-source web app built on a structured analysis of that landscape across more than 50 engineering organizations. Alongside the tool, we share what the dataset reveals about how DevEx is measured today and where the gaps lie. Whether you are new to DevEx measurement and looking for a first meaningful set of metrics, or a seasoned practitioner assessing whether an existing collection has the right breadth and depth, Compass helps teams navigate the fragmented landscape and identify measures that actually matter for their context and goals.

Too Many DevEx Metrics, Too Little Guidance - ACM Queue As AI-augmented development becomes standard practice, engineering leaders face mounting pressure to demonstrate its impact. Yet most organizations are measuring AI adoption and output while the effect on developer experience (DevEx) remains largely unknown. The right metrics can close that gap, sur...

Want your school to be the top-listed School/college in New York?
Click here to claim your Sponsored Listing.

Address


2 Penn Plz
New York, NY
10001

Opening Hours

Monday 9am - 5pm
Tuesday 9am - 5pm
Wednesday 9am - 5pm
Thursday 9am - 5pm
Friday 9am - 5pm

Alerts

Be the first to know and let us send you an email when ACM Queue posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Subscribe

We will notify you when anything happens in New York.