On-premises · Australian-made AI

Matilda, on your GPUs

The models we train in Melbourne, trained further for your domain and delivered into your data centre, so you take the weights, the serving stack, the whole API, or the entire product. It runs behind your own perimeter, on hardware you already have or on a cluster we build for you, air-gapped if the work demands it.

Melbourne
Where the models are trained
Ours
The weights are ours to hand over
Your domain
What we continue training it on
Five
Levels of hand-over
Deployment levels

Take the stack as far down as you want it

Each level contains the ones before it, and most teams start at the API and move down when their accreditation asks them to, so nobody has to take the whole thing to get started. Plenty of deployments never touch level one, because the racks are already there.

Level 01 · the cluster

The machine as well, built to the same blueprint

Most organisations that want a model inside the perimeter do not yet have anywhere to put it. We design, build, and validate the cluster too, to the same blueprint we run for MC-1 and MC-2, so what lands in your data centre is the architecture our own models are trained and served on rather than a reference design nobody has had to live with.

Both platforms, both in production

Accelerators
  • NVIDIA

    CUDA

    We run NVIDIA accelerators in production for inference, and where NVIDIA is what your estate is standardised on, that is what we build for you. It is the most mature serving path in the industry and we treat it as a default rather than an exception.

  • AMD

    ROCm

    Our own production serving fleet runs on AMD, and getting vLLM to perform properly on ROCm is work most teams start and then abandon. We have done it at cluster scale, which is what makes an AMD estate a first-class option here rather than a compromise.

The two behave differently under training and under serving, and part of the design work is placing each workload where its platform is strongest. Whichever one is already in your racks is the one we support.

How it is handed over

Built twice, on purpose

The environment is defined in playbooks rather than in somebody's shell history, and the way we prove that is by building it, wiping it, and building it again from the same playbooks.

The first pass brings the cluster to spec and exposes whatever automation is still missing. Then the machines are flashed back to bare metal and the same playbooks run again, which is the pass that proves there is no snowflake state left in the building. That is the difference between a cluster that was installed once and a cluster you can rebuild on a bad day.

Gates before handover

  • Every node joined and ready
  • Networking and cluster DNS healthy
  • Storage pools and classes present
  • GPUs allocatable on every accelerator node
  • Database, cache, and serving suites passing
  • Saturation and failover exercised under load
Model customisation

Your domain, trained into the weights

Retrieval and adapters attach context to a general model at its edges, and they hold up until the task gets genuinely hard. We continue training Matilda itself on your domain and your task, working inside the network rather than around it, so the capability ends up in the weights rather than in the prompt.

Where the change lands

How deep
  • PromptingThe instruction you send
  • RetrievalWhat sits in the context window
  • AdaptersA thin layer over attention
  • Continued trainingThe weights, and how the model reasons

Each has its place, and retrieval will always matter for facts that change by the hour, but a task the model has to be genuinely good at belongs in the weights.

Why it is hard to buy elsewhere

We trained Matilda, so we know what each stage does

Matilda exists because we trained it ourselves from a base model up, which means we know what each stage of training changes and what it costs to run. Post-training is not one recipe, and the strategy that suits a bank is not the one that suits a defence workload, so choosing it is most of the skill.

The work also needs a cluster, and for material that cannot leave the country, it needs that cluster and the people running it to be here. A model trained specifically for adversarial security review is a different piece of work from one trained to read case notes, and both are the kind of thing we would take on.

Containment

It runs on your hardware, not ours

The deployment runs on your hardware, inside your network, under your change process, and the claims below are ones your security team can check for themselves rather than take on trust.

Built for work that cannot leave the building

04 sectors
Defence and national security
Isolated networks and long accreditation cycles
Banking and financial services
Customer data under obligations that name the boundary
Government and critical infrastructure
Onshore residency and a control set your assessors know
Health and medical research
Patient and participant data that stays in the institution
Australian-made AI

Australian-made, all the way down

It matters who trained the model, who built the weights, and where your traffic goes, and those questions have short answers here, which is the whole reason this offering can exist.

How we build in Australia

Who trained the model
Maincode, in Melbourne, on hardware we own and operate
Who built the weights
We train them ourselves, which is what makes handing them over possible
Where the company sits
Australian, incorporated and headquartered in Melbourne
Where your traffic goes
It stays on your hardware, so it is never ours to hold
Whose hardware trained it
Ours, the MC-2 AI Factory in a Melbourne data centre
Who you call when it breaks
A named engineer in Melbourne, in your timezone
Accreditation

Built for the assessment

A deployment inside your perimeter is your system, so the authority to operate is yours to hold, and our job is to supply what your assessors ask us for while they work through it.

  • Architecture and data-flow documentation for the deployed system
  • Control mapping for the components we ship
  • Software bill of materials and signed release artefacts
  • Configuration baselines for the environment as installed
  • An engineer on the call with your assessors
How it lands

From the first call to the first token

A Matilda deployment is an infrastructure project rather than a signup, and it moves at the speed of your change board and your assessors, so we plan it that way from the start.

  1. 01

    Scope

    What the work is, what it is classified at, and what silicon you already have, and if it turns out you do not need on-premises at all, this is the call where we say so.

  2. 02

    Design

    A reference architecture for your environment, covering model sizes, node counts, network topology, identity, and where the audit stream lands, which your security team reviews before anything is ordered.

  3. 03

    Install

    We deploy alongside your platform team, on your change process, and whether the site is connected, isolated, or fully air-gapped, the artefacts are the same, with only the delivery differing.

  4. 04

    Accredit

    We supply the evidence pack and sit with your assessors while they work through it, and while you hold the authority to operate, our job is to keep the questions cheap to answer.

  5. 05

    Operate

    Model updates on your schedule, benchmarked before they land, with a rollback path, and a named engineer to escalate to rather than a ticket queue.

Before you bring it inside

The questions security and procurement ask, answered before the first call, with the rest in the docs.

Bring the model to the work

Tell us what the work is and what you are running on, and we will tell you whether Matilda fits and at which level, including when the answer is that you do not need any of this.