Explainer · AI in Transportation
AI in transit: where it's real and where it's a demo
Transit agencies are running machine learning in production today, but not in the places the marketing points to. Here is what the working deployments have in common, and why the ones that stall usually stall on data rather than on models.
Ask whether artificial intelligence has arrived in public transit and you get two answers, both wrong. One says the technology is transforming how agencies run service. The other says it is vendor theater. The accurate answer is narrower and more useful: machine learning is running in production at US transit agencies right now, it is doing real work, and almost none of that work is happening where the marketing points.
The deployments that hold up share a shape. They sit on a data stream the agency was already collecting for some other reason, they answer one narrow question instead of managing a whole process, and they hand the answer to a person who makes the final call. Where that shape holds, the results are measurable and faintly boring. Where it does not, an agency tends to end up with a pilot, a press release, and no second phase.
What agencies say they are actually running
The most useful recent survey of the field is APTA’s, published in May 2026 as Artificial Intelligence and Machine Learning in Public Transit: A Primer, alongside four companion guidance briefs. It draws on responses from 32 member agencies, staff interviews, and a review of documented deployments, and it sorts what it found into eight functional areas: back office, operations, customer support, maintenance, safety and security, customer analytics, planning, and fares and ticketing.
The finding worth sitting with is where the activity is concentrated. Customer support and customer analytics are the most common areas of current deployment. Count planned deployments too, and back office and operations move to the top, at 50 percent and 47 percent of respondents respectively reporting current or planned use. So the center of gravity is call handling, scheduling support, document processing, and analytics. Autonomous vehicles do not appear, and neither does anything that touches a safety-critical decision without a person in the loop.
That distribution is not timidity. It tracks where agencies already have the thing a model needs, which is a high-volume record of what happened, labeled well enough to learn from.
The primer names five agency deployments, and the list is worth reading as a group because of how ordinary it is.
| Agency | Where it runs | What it does |
|---|---|---|
| MTA (New York) | Maintenance | Predicts component failure on the bus fleet from maintenance history |
| AC Transit (California) | Safety and security | Image recognition on bus-mounted cameras to detect vehicles blocking bus lanes and stops |
| Riverside Transit Agency (California) | Customer support | Pushes real-time detour and disruption updates out across rider channels automatically |
| CapMetro (Texas) | Operations | Virtual agent that books paratransit trips, leaving staff the complicated calls |
| Prairie Hills Transit (South Dakota) | Operations | Automated dispatch, replacing a handwritten scheduling process |
Not one of them decides anything a rider would notice without a person in the loop, and the last is a rural operator in a state with fewer than a million people. This is what deployed transit AI looks like in 2026: narrow, internal, and aimed at work that was already being done by hand.
The case that works
Automated camera enforcement is the example worth studying, because one agency published a before-and-after on the same corridor with the same review process at the end of it.
AC Transit runs camera enforcement on its Tempo corridor in the East Bay, a 9.5-mile line with 100 buses carrying forward-facing cameras that watch for vehicles stopped in bus lanes and at bus stops. In June 2024 the agency replaced its legacy setup, which required manual camera activation, with software that recognizes lane markings, bus stop dimensions, and vehicle sizes. Who decides did not change. Trained law enforcement personnel, and not the software, review each evidence package and determine whether a citation is issued. The software’s job is to decide what gets looked at.

The raw counts are striking but the ratio is the real finding. Under the legacy system, 22 citations came out of 879 evidence packages, so about one package in forty was worth acting on. Under the AI system, 787 citations came out of 1,102 packages, roughly seven in ten. AC Transit characterized the change as a 34.4-fold increase in citation efficiency. Because the two observation windows are not the same length, the per-package rate is the more honest way to read it, and it points at the same conclusion.
Neither figure counts how many violations happened on that corridor. The legacy number especially does not, since a system that depended on an operator activating the camera only ever captured a fraction of what occurred. What the comparison measures is the quality of the queue, which is the thing a reviewer’s day is actually made of.
It is worth being careful about why the queue improved, because two things changed at the same time and the published figures cannot separate them. The cameras stopped needing an operator to trigger them, and the detection software got better at recognizing what it was looking at. The 1-in-40 conversion rate under the old system is itself a clue that the earlier problem may have been evidence quality rather than detection: a package fails review when the plate is unreadable or the clip misses the moment, not only when nothing happened. On that reading the win is careful video engineering as much as machine learning. Both stories land in the same place, which is the useful part. The system got better at producing something a person could act on, and that is what changed the output.
The same pattern is now running at scale in New York, where the MTA’s Automated Camera Enforcement program reached 67 bus routes as of July 2026, with more than 1,900 equipped buses covering roughly 810 route miles. The MTA reports a 20 percent reduction in collisions on ACE routes and speed gains approaching 30 percent on some segments. The agency’s own framing on that second number deserves to be carried over intact: it credits the gains to ACE combined with dedicated bus lanes and street upgrades. The redesign and the camera are doing the work together, and the caveat tends to fall away as the figure gets repeated.
The other headline ACE result needs the same care, and it is a useful lesson in reading agency numbers. The program is usually described as having cut blocked bus stops by 40 percent. What the underlying measurement shows is a 40 percent drop in the number of drivers ticketed for blocking them across the routes ACE has covered since June 2024. On the Bx36 the count fell from 4,905 tickets in September 2024 to 754 in November 2025. Deterrence is the most plausible reading of a decline that steep, and the drop is almost certainly real. It is still worth noticing that the camera is both the enforcement mechanism and the measuring instrument, so the system is grading itself with a device it controls. That is not a reason to throw the number out. It is a reason to want the denominator published next to it.
What the working deployments have in common
Set these cases side by side and four things recur. This is a description read off the same handful of deployments it describes, so treat it as a set of better questions to ask rather than a test with predictive power. The striking part is that none of the four is about model architecture.
| Condition | What it means in practice | Why it matters |
|---|---|---|
| An existing data stream | The signal is already collected for some other purpose: vehicle location reports, farebox records, call transcripts, camera video, maintenance work orders | A pilot that must first build a sensor network rarely reaches phase two. See the correction below, which matters more than the rule |
| A narrow, well-posed question | “Is this vehicle stopped in a bus lane” or “which component is trending toward failure,” not “optimize operations” | A narrow question has a right answer you can check. A broad mandate has no test that can be failed |
| A person at the decision point | The model produces a candidate; a reviewer, dispatcher, or mechanic decides | Keeps errors recoverable, keeps accountability with staff, and makes the system deployable under existing policy |
| A measurable baseline | The agency knows what the previous process produced, on the same corridor or fleet | Without a before, there is no after. This is the condition most pilots quietly skip |
The first condition needs a correction, and the correction is the more useful half. AC Transit’s upgrade rode on cameras that were already mounted on the buses, which is what let it prove something cheaply. What happened next looks nothing like that. The MTA now runs more than 1,900 camera-equipped buses, and New York is funding 200 additional stationary cameras on top of them. That is a new sensor network, built at scale, for exactly this purpose. Applied too strictly, the rule would have predicted camera enforcement never gets past the pilot corridor. The honest version is a sequence rather than a rule: an existing stream is what lets a narrow win get proven cheaply, and a proven narrow win is what unlocks the budget for sensors. Proposals that ask for the sensor network first, before anything has been shown, are the ones that stall.
The last row is the one that separates a program from an anecdote. AC Transit could show its result precisely because it had run the old system on the same 9.5 miles and knew what it produced. Most agency AI announcements have no comparable baseline, which is not evidence the tool failed, but it does mean nobody can say whether it worked.
APTA’s guidance briefs land in roughly the same place, addressing tool and infrastructure needs, policy and governance, agency readiness and staff capacity, and implementation. Read as a list of what agencies struggle with, that is an honest one. The hard parts are procurement, data plumbing, and governance.
Where it is a demo
Set that against the weapons-detection pilot run in the New York subway in the summer of 2024. Over roughly a month across 20 stations, the scanners performed 2,749 scans. They found zero guns, recovered 12 knives, and produced 118 false positives. The Legal Aid Society called it objectively a failure, and the NYPD said it was still evaluating the results, having entered into no contract with the vendor.
Two caveats before drawing anything from it. This was an NYPD and City Hall program that happened to run inside transit stations, not a transit agency’s own deployment, so no agency data team selected it or has to live with the result. And the conditions above were written with cases like this one in view, so announcing that it fails them proves nothing.
What it does illustrate is a failure mode worth naming on its own terms. When the thing being searched for is vanishingly rare in the population being scanned, a small false-positive rate becomes a large absolute number of stopped riders.

That is not a bad model score. It is an arithmetic property of screening for rare events, and no improvement to the detector removes it. The consequence also landed on a rider immediately rather than on a reviewer working through a queue, which is the structural difference between this and the camera systems.
Worth stating plainly: no well-documented failure of a transit agency’s own AI deployment turns up in the APTA survey or in the trade record. That absence is not reassuring. Agencies publish launches. Very few publish post-mortems, and an agency that quietly stops using a tool has no reason to announce it.
The same evidence problem, in a milder form, runs through the current vendor wave. Optibus launched what it calls the first AI agent purpose-built for public transportation on June 17, 2026, covering planning, scheduling, dispatch, and live operations, and bundled into existing customer licenses at no extra cost. The capability list is plausible and the problems it names, driver shortages and zero-emission fleet transitions among them, are real pressures on scheduling teams.
What the launch does not carry is a named agency running it or a single quantified result. That is worth stating carefully. Absence of published evidence is not evidence of absence, and a product a few weeks old has not had time to generate a season of operating data. Call it unproven and watch for whether named deployments with baselines follow within the year. One thing the bundling does not tell you is whether the thing works: shipping the agent inside existing licenses at no charge is a way to get it adopted, and adoption is not a result.
The data layer decides, again
The pattern underneath all of this will be familiar to anyone who has followed how real-time arrival predictions are actually produced. The machine learning step is real and it does measurably better than the naive method on clean, high-volume data. It also sits on top of an unglamorous pile of prerequisites, and it fails in the same ways that pile fails.
Predictive maintenance shows it plainly. The model in a failure-prediction system is not exotic. What makes one possible at all is years of work orders, parts consumption records, and telematics history, recorded consistently enough that a component’s slide toward failure is visible in the data. An agency without that history cannot buy the result, because the result lives in the history rather than in the software. APTA’s primer does put a number on one such deployment, crediting the MTA with a 75 percent gain in maintenance productivity and a 24 percent cut in material costs across its bus fleet. Hold that figure loosely. It is agency-reported, with no published baseline, period, or method behind it, and it is the kind of number that gets repeated precisely because it is striking. The claim that survives without it is the structural one: the recorded history is the asset, and the model is the cheap part.
The same logic explains why ghost buses do not yield to better prediction models. A vehicle assigned to a trip it is not running produces confident, well-formed, wrong output no matter how good the estimator is. Feed the model bad trip assignments and it will learn to reproduce them.
This is the practical test to apply to any transit AI proposal. Ask what stream it learns from, whether the agency already has that stream, how many years deep it goes, and who checks the output. An answer that leads with model capability and gets vague about data provenance is describing a demo.
What to watch
A few things over the next year will say more than any product launch. Watch whether agencies publish baselines alongside deployments, as AC Transit did. That one practice would let the field tell tools that work apart from tools that were merely bought, and almost nobody does it. Watch procurement language too, specifically whether contracts begin specifying data ownership and model performance in terms someone could measure later. That gap is documented. A review of 13 federal AI acquisitions published by the GAO in April 2026 found officials at every agency it examined naming requirements definition and contract terms as a core difficulty, and found FEMA unable to share model outputs with state partners because the data rights were never secured at award. Those are federal agencies rather than transit ones, but the contracts are written the same way, and APTA’s governance and readiness briefs exist because the same gap runs through transit.
Camera enforcement is about to get a large natural experiment. The Next Stop plan announced for New York on July 8, 2026 names 50 priority corridors, expands ACE to 25 additional routes in each of 2026 and 2027, and adds 200 stationary bus lane cameras by 2027, alongside new protected lanes and rapid bus corridors. More enforcement and more street redesign are landing on the same corridors at the same time. That is good for riders and unhelpful for anyone hoping to attribute the resulting speed gains to one or the other. If the cameras deserve credit, this is the moment to instrument for it.
Then there is the vendor agent wave now arriving. If named agencies are running these agents on live scheduling and dispatch a year from now, with published operating results, that is a genuine expansion of what the technology does in transit. If the launches are still describing capabilities and citing no operators, that will be its own answer. This is a fast-moving area, so treat the specifics here as a snapshot of mid-2026.
The framing that survives all of it is simple enough. AI in transit is neither arriving nor overhyped as a category, because it is not a category. It is a set of narrow tools that work well when they sit on data an agency already has and report to a person who can overrule them. The agencies getting value from it are mostly the ones that did the boring data work first, which means the question to ask about any proposal is not how capable the model is. It is what it is standing on.
Common questions
- Are transit agencies actually using AI today?
- Yes, though mostly in the back office rather than in service delivery. APTA's 2026 survey of member agencies found customer support and customer analytics are the most common current deployments, with back office and operations topping the list once planned deployments are counted.
- What kind of transit AI works best in practice?
- Deployments that sit on a data stream the agency already collects, answer one narrow question, and hand the result to a person who makes the final call. Automated camera enforcement is the clearest example: the software flags candidate violations, and trained staff decide whether a citation is issued.
- Why do transit AI pilots fail?
- Usually because of the data underneath rather than the model on top. Proving a first win cheaply depends on a stream the agency already collects and on a baseline showing what the old process produced. Proposals that ask for a new sensor network before anything has been demonstrated are the ones that tend to stall.