Research portfolio

More capability.
Clearer evidence.

Keep provides a shared setting for investigating how autonomous systems learn, act, coordinate, and recover. The questions below describe areas of investigation—not claims that each problem has been solved or that the underlying techniques are novel.

01 / Learning

Learning, memory,
and changing evidence

When does retained experience improve performance on genuinely new work? How should an agent handle corrections, withdrawals, and conflicting information without retaining invalid guidance or discarding useful knowledge unnecessarily?

Keep provides source-linked memory and adaptation mechanisms for investigating these questions. Broad learning transfer remains to be established through comparative evaluation.

02 / Skills

Useful and controlled
skill acquisition

How can an agent acquire or reconstruct a procedure, establish that it actually helps, and keep its use within appropriate limits? When should a candidate be promoted, revised, or retired?

Pattern extraction, structural validation, task performance, and permission to execute are different questions. Research must evaluate them separately rather than treat a “validated” label as a complete answer.

03 / Execution

Native execution
and verification

Which combinations of candidate generation, tests, acceptance policies, and operator review improve correctly completed repository work? What additional cost or human effort do those improvements require?

Keep's native coding and acceptance mechanisms provide an implementation base. Generated tests remain proposals to evaluate; passing a test suite is not proof of unrestricted correctness.

04 / Coordination

Coordination, authority,
and recovery

How should interacting operations share resources and authority? What should happen when an effect becomes uncertain, or its original permission or supporting information changes before recovery completes?

Controlled failure cases allow comparison of actual effects, resource obligations, and subsequent useful work. Complete behavior across all execution paths remains an open qualification task.

05 / Human control

Human control, privacy,
and isolation

How can people delegate meaningful work without transferring unnecessary authority or information? Which protections actually hold under a particular arrangement of credentials, processes, tools, and providers?

Both useful task completion and the cost of oversight matter. Configured controls must be distinguished from protections observed in operation.

06 / Evidence

Evidence and
reusable evaluation

What can an evaluator establish about authorization, execution, recovery, and missing observations? Which tools could help examine systems other than Keep?

The aim is to make relevant outcomes and limitations inspectable. Hash-linked records can support integrity checks; they do not, by themselves, prove truth, complete observation, or independent oversight.

The evaluation approach

How progress is evaluated

We distinguish implemented mechanisms, controlled demonstrations, and comparative results. Evaluation should include competent alternatives, held-out work where appropriate, actual outcomes, resource costs, and cases in which the proposed approach fails. A useful negative result can be a reason to simplify the system or change direction.

Resources that expand
the questions we can test

Additional model access, isolated infrastructure, controlled local models, and specialist collaboration can enable broader comparisons. Weight adaptation and accelerator-dependent work require their own implementation and qualification; access to hardware alone is not a demonstrated research result.

An open conversation

Research support
and collaboration

Red Rook AI is exploring support for well-defined research and reusable technical outputs across this portfolio. Potential contributions include research funding, model or compute access, independent evaluation, and specialist collaboration.

Discuss a research opportunity