Long-form articles on the parts of AI adoption that decide whether a project creates value. Every claim is traced to a published study, regulation or standard — not to vendor material.
8 articles · 60 sources cited · peer-reviewed research, regulation and public standards
Only about a tenth of companies report significant financial benefit from AI, and practitioners believe AI projects fail more often than conventional IT projects. The research literature is remarkably consistent about why: the decisive mistakes are made before any model is trained.
The published evidence points at problem framing, data ownership and organisational learning — not model accuracy — as the dominant causes of failure.
Putting a human in the loop is treated as a guarantee of safety. The empirical record says otherwise: reviewers miss errors, over-rely on confident systems, and abandon accurate models after a single visible mistake. Effective oversight has to be engineered.
A meta-analysis of 106 studies found human–AI combinations on average performed worse than the better of human or AI alone. Oversight only helps when it is deliberately designed.
Most AI business cases are built on vendor claims. There is no need for that: several large field experiments now report clean effect sizes, and they tell a consistent story about who gains, how much, and where the estimate collapses.
Randomised studies give defensible ranges — 14 percent more support issues resolved per hour, around 40 percent less time on writing tasks — but the gains concentrate among less experienced staff.
Compliance work done before a pilot is cheaper than compliance work done after it. For a Swedish company adopting language models, four questions decide the workload — and all four have citable answers in the regulation itself.
The AI Act applies in phases: prohibitions since February 2025, general-purpose model obligations since August 2025, most high-risk duties from August 2026.
A demo runs on a curated document. A product runs on a scanned PDF, a ticket written in three languages, and a Monday morning. The engineering literature has described that gap precisely for a decade.
The model is a small share of a production ML system. Everything around it — glue code, data dependencies, monitoring, evaluation — is where the cost lives.
The interface decides whether a fallible system is useful or dangerous. Two decades of HCI research, including several controlled studies at CHI and FAccT, show that the obvious design moves — add an explanation, show a score — do not do what teams assume.
Calibrated confidence improved trust calibration in controlled studies; explanations alone often increased reliance without making it appropriate.
A smaller company does not need a governance department. It needs a small number of living documents that map onto recognised frameworks, so that a customer, an auditor or a regulator can see the same structure they expect from a large organisation.
Four artefacts carry most of the value: an inventory with risk levels, documented models and datasets, a written oversight description, and an incident log.
Deployment is not adoption. Decades of information systems research identify what actually drives use, and recent studies of AI resistance show that the people you expect to champion the tool are often the ones who resist it.
Adoption research points at performance expectancy, effort expectancy, social influence and facilitating conditions — and your strongest performers may resist most.