Applications and agents
Connect interfaces, backend services, records and tools. Define supported actions and approval steps, then test the result in the application where the work happens.
Technical capabilities
Model evaluation, application engineering and controlled deployment, supported by a research and delivery practice that records what was tried, what it cost and how it performed.
What we bring
Connect interfaces, backend services, records and tools. Define supported actions and approval steps, then test the result in the application where the work happens.
Compare models and workflows against retained tasks and acceptance criteria. Measure useful results, failures, response time and the effort needed to review and repair.
Work with managed services, dedicated Australian hosts, client cloud or on-premises infrastructure. Establish the data paths, configuration and operating responsibilities.
Model selection
We evaluate open-weight models alongside managed offerings. An implementation task, an independent review and a coordination task each place different demands on a model. We test those roles explicitly.
Run representative tasks against retained source versions and acceptance criteria. Check working behaviour, defects and repair rounds, including the result after an action in connected software.
Assess missed defects, false alarms and the usefulness of findings. Review quality needs its own evidence; a model’s ability to produce an implementation does not establish its ability to assess one.
Examine task division, coordination and successful integration. Capture runtime, token usage, failures and human intervention across the workflow.
Automated evaluation lets us assess new models promptly using repeatable tasks. In a September 2026 evaluation, ten comparative builds across three task types ran within a day. Adoption follows the results and the requirements of the environment where the model will be used.
Open-weight models
Open-weight models provide options beyond a single managed provider. On controlled compute, we can retain the weights, tokenizer, runtime, quantisation and generation settings alongside the workflow. Proposed replacements can be tested before they change a working system.
We assess task quality, licence, data boundaries, infrastructure and support requirements together. A lower-cost model may suit one role while another needs a different approach. Running costs include cached and uncached input, generated output, retries, repairs and review.
How we assess operating economicsResearch and delivery evidence / Snapshot: 14 September 2026
Across 111,624 metered requests in our instrumented research and delivery environment, recorded 17 August–14 September 2026.
The retained records expose expensive stages, weak reviews and failed attempts. We use them to improve model selection and the way engineering work is executed and checked.
Approximately 91% of these tokens were cache reads: reused context. Routed usage includes managed endpoints and is not a measure of local GPU output. Components are rounded.
A separate registry records 484 benchmark and evaluation runs and 587 engineering workflows from 9 July–14 September 2026, Sydney dates.
These are deduplicated launches, including experiments and unsuccessful attempts, rather than accepted deliverables. The registry and token ledger cover different periods. This is a dated record of recent activity, alongside a delivery history extending over a decade.
Source: retained usage ledger and launch registry reported in Form From’s Edition 07 technical overview.
A measured comparison
In a repair-task comparison recorded 1–3 September 2026, an open-model builder used A$2.75 in model-use cost versus A$20.05 for the reference run. It took 65 minutes versus 36 minutes. The same blind reviewer accepted the core repair in both outputs and requested further corrections.
That is a task-specific cost and time trade-off to investigate. The comparison included one run per builder, excluded host and review time, and had no post-merge behaviour test. It does not establish equivalent quality across models.
For a client decision, we agree representative work, repeat the comparison and include the costs and acceptance checks that matter to the organisation.
Delivery practice
Our AI-assisted engineering method scopes work into reviewable changes, separates implementation from review and checks integration against the current codebase. Human decisions remain explicit at acceptance and release.
We can work through your repositories, issue tracking and CI/CD. Our internal tools support research and execution where useful; the engagement is shaped around your delivery environment.
At ResetData, delivery included practical AI-development sessions with engineers on their own product. That connected tool use to task specification, review and release, helping the team continue after handover.
Read the ResetData case studyFrom architecture to operation
Our experience includes privacy processing, permission-aware workflows, repeatable infrastructure provisioning with Terraform and Ansible, and records linking source information, review decisions and resulting actions.
Australian deployment is available. Its scope includes model processing, application data, telemetry, backups and connected services. We work with your technical and governance owners to establish the actual requirements and test the relevant controls.
Let’s make it work
Tell us what you want to improve, which systems are involved and what is still uncertain. We’ll help establish the next useful step.
Start a conversation