Applied AI research / Agenda in development

Dependable
AI Systems.

Research at the point where human judgment, machine-generated evidence and autonomous action meet. This is an active direction, not a claim of completed findings.

01 Core research question

Core research question

How can humans safely delegate increasingly complex decisions and actions to autonomous AI systems?

From interpretation to action

Three connected
levels of dependence.

Reliability is not only a property of a model. It emerges through the relationship between what a person understands, what a machine can substantiate and what a complete system is allowed to do.

01

HumanInterpretation & reliance

How communication shapes perceived authority, trust and the decision to accept, verify, challenge or reject an AI-generated answer.

  • Trust
  • Authority
  • Human–AI interaction
  • Communication of uncertainty
  • Decision-making
  • Verification behavior
02

MachineEvidence & uncertainty

How a model’s claims can be grounded, calibrated and traced—and when a system should retrieve, clarify, abstain or escalate.

  • Evidence
  • Grounding
  • Uncertainty
  • Calibration
  • Provenance
  • Factuality & abstention
03

SystemAutonomous action

How agents plan, use tools and recover across long-running tasks—and how reliability, security and safe action can be evaluated at system level.

  • Autonomous agents
  • Planning
  • Tool use
  • Agent evaluation
  • Security & recovery
  • Safe autonomous action

03 Communication → computation

The bridge is not artificial.
It is the research problem.

Interactive communication asks how a system constructs authority and how a person interprets certainty. Dependable AI asks how that authority and certainty can be justified technically.

Questions about conversational evidence and user trust lead directly to grounding, provenance, calibration and verification. As systems become more autonomous, those same questions become operational: when should a system act, seek more evidence, ask for help or refuse?

Work in progress

Research direction

Questions before
findings.

  1. How does the communication of uncertainty affect human reliance on LLM-generated answers?
  2. Does displaying evidence improve a user’s ability to detect incorrect AI responses?
  3. How do AI systems construct perceived authority in conversational interfaces?
  4. When should an autonomous system answer, retrieve more evidence, ask for clarification or abstain?
  5. How can agent reliability be measured over long-horizon tasks?
  6. How should confidence, evidence and uncertainty be communicated to users?

Public record

Research output,
when it exists.

Research in progress. Working papers and publications will appear here.

No publications listed

Reproducible work

Evidence should leave
an artifact.

Datasets, evaluation harnesses, simulations, reproducible experiments and research prototypes will be kept separate from general engineering projects.

Research artifacts will appear as the agenda moves into experiments.

No artifacts listed