What are the implications?
If this system is wrong, who is affected and how badly? An AI drafting internal notes and an AI talking to your customers carry completely different consequences, and they should not be engineered to the same standard.
The discipline
The tools got easy enough that anyone can build something impressive in an afternoon. The hard part never moved. It is still reliability, security, data handling, and knowing when something has gone wrong.
In short: AI systems engineering means treating an AI deployment as a system to be engineered rather than a product to be bought. It covers reliability, security, data handling, backups, failure detection and operational ownership. Building something that works takes days; making it hold up in production takes months. AI4SMB applies this to Australian small and medium business, built on a managed services provider with 32 years of running business systems.
The shift
That sentence is not a criticism of anyone. It is the defining fact of this moment. The barrier to producing something that works has collapsed, and the barrier to producing something that keeps working has not moved at all. So businesses now run on systems that were assembled quickly by capable people who were never taught what a system needs, and mostly they are fine, right up until they are not. The skills that close that gap are not new or exotic. They are the ordinary disciplines of running business infrastructure, applied to a new kind of system.
The questions
None of these are AI questions. They are the questions you would ask of a phone system, a file server or a backup regime, and they work here for the same reason they worked there.
If this system is wrong, who is affected and how badly? An AI drafting internal notes and an AI talking to your customers carry completely different consequences, and they should not be engineered to the same standard.
Not "does it work in the demo". What is its behaviour on the awkward input, the unexpected question, the day the vendor changes something upstream. Reliability is measured at the edges, not in the middle.
What can it reach, who can talk to it, and what could someone make it do that you did not intend. AI systems are reached through language, which means the attack surface includes anything the system reads.
Which service, which country, retained for how long, used to train what. Under the Privacy Act the obligation to know sits with you, not the vendor, and the answers should be in writing.
The prompts, the configuration, the accumulated context, the integrations. Plenty of businesses have an AI system that would be genuinely difficult to reconstruct and have never thought of it as something to back up.
The one people never ask, and the one worth the most. AI does not usually crash. It produces a confident wrong answer and carries on, so the failure is silent unless somebody engineered a way to see it. Worse, so is the failure of whatever was meant to be watching.
The deepest version of that last question
Almost every monitoring setup is built to speak only when something is wrong. It is a sensible instinct and it contains a trap, because it makes silence ambiguous. When nothing arrives, it means one of two things: everything is fine, or the thing that would have told you is itself broken. You cannot tell those apart from the outside, and you will assume the comfortable one, because that is what people do. This is how service failures go unnoticed for days inside otherwise well-run automated systems. Nobody ignored an alert. No alert was ever sent.
The fix is to invert the logic. Instead of asking the system to speak up when something breaks, you require it to prove it is alive on a schedule, and the alarm fires when the proof stops arriving. Silence stops being a state you have to interpret and becomes the failure condition itself. It is an old idea, it has an unlovely name, and it is the single most valuable thing most small and medium businesses are missing from their monitoring.
Two things have to be true for it to work. It has to run somewhere else, because a watchdog living inside the system it watches dies with it and tells you nothing. And it has to check the business outcome rather than the service, because a server can be up, green and doing nothing useful at all. "Is the service running" is a weak question. "Have any tickets moved in the last hour, is data still flowing, did the AI engine answer anything today" are the questions that actually correspond to your business working.
The honest arithmetic
This is the part that never appears in anyone's marketing, so here it is plainly. Standing up something genuinely useful with AI takes about three days, and the result will impress you, because it is impressive. What follows is access control, failure handling, monitoring, data-retention decisions, testing against how it actually breaks rather than how you hope it works, and documenting it so it survives the person who built it. That is months. None of it is visible and none of it feels urgent. Nobody thanks you for the outage that did not happen or the breach that did not occur. But the gap between the three days and the three months is the entire difference between a demo and a system, and it is the only part worth paying anyone for.
The uncomfortable part
In 2022 the records of roughly ten million Australians were taken from Optus. The cause, as publicly reported, was not a sophisticated attack. An interface that returned customer details had been published to the internet with no password on it at all. Worse, the customer records were numbered in sequence, so once you found the door you did not need to be clever. You counted upwards. Nothing limited how fast you could ask, and it stayed open for up to three months.
Put plainly: the front door had no lock, and the filing cabinets were numbered in order. Nobody had to break in and nobody had to search.
That is what a real breach usually looks like, and it is why the arrival of AI does not change the security conversation as much as people expect. AI adds genuinely new attack surfaces, and prompt injection is a real one that we have measured on our own systems. But the rule it breaks is the oldest rule in the trade: do not trust input from outside your boundary. The foundations of good IT have not changed. What changed is how many people are now building on top of them without knowing they are there.
Proof rather than assertion
We ran six open-weight AI models through a diagnostic harness, then hid a malicious instruction inside the data they were reading. Two of the finalists recommended destroying a machine's ability to recover. The fix was not a better model, it was two sentences added to the instructions, after which obedience to the attack went to zero and answer quality went up. We published the whole thing, including the part where our own assumption was wrong. Read it here. The same discipline runs on this website: it is built and maintained by AI, and the results are published at /experiment whether they flatter us or not.
Where it goes
Where AI pays back, what it would touch, and what would need engineering before it could be relied on. From $800 ex GST.
The ongoing half. Monitoring, maintenance, improvement and an accountable support path. Engineering is why it works; this is what you buy.
For the data that genuinely cannot leave. Run on infrastructure you control, with the operational cost stated honestly.
Book an audit call. We will tell you what your AI actually touches, how it would fail, and whether it is built well enough to rely on.