What is good enough?
AI is more than just prompt engineering
How do we manage costs with AI? How do we build judgment with AI? How do we think about buy vs. build?
I did a talk for the Singapore Venture & Private Capital Association today (deck below).
I was blunt. Aside from a brief sojourn with PEVC during my days heading investment risk management in MAS, this is not my world.
And I tried to convince the audience about the need to stop thinking of AI as this monolithic thing. Or ditching tools that work for the lure of LLMs. Or falling prey to the mythology of an AI that schemes and manipulates because that hurts more it helps.
And that learning prompt engineering or design is a farce. It might seem fascinating when some AI 'expert' tries to teach you the intricacies of how a prompt should be constructed. And all the nonsense about how "X is the best way to write the prompt"; "Y is the best length for a prompt"; "Be careful about a Z tone when you prompt AI".
If learning AI is about learning these 'truths', then such nonsense needs to be evaluated and tested on its efficacy properly. Not based on a single demonstration. Also, remember that there is a high chance it may not hold for the next model.
To be honest, I am surprised prompt engineering is still a thing. Many of the ways of working with AI agents these days have made prompting a lot less important. Thinking about it as a system works better.
But I think the key questions on the minds of most in the audience was not the above. But about costs, judgment, buy vs. build. I think that these are problems on the surface.
At the core, the question we should be asking ourselves is this - do we even know what is good enough for our jobs?
And I know this sounds boring. But at the end of the day, every one of those surface questions is about the same question. Cost is really "is it good enough to be worth what I pay for it?" Judgment is "do I know what is good enough when it tells me something?" Buy versus build is "is what I can buy good enough, or do I have to make my own?"
You cannot answer a single one until you can say what "good enough" means for your task, and show that the AI actually clears that bar you know best. That is the core task. Not writing prompts. Knowing what good looks like for the job in front of you, and being able to prove that AI meets it.
That is all evaluation and testing really is. An unglamorous name for building a set of real examples from your own work, deciding what a good answer looks like, and running the AI against it. No demo. No vibes. A bad AI system serves no one, and the only way you find out whether yours is bad is to test it against what your customers and what your work actually needs.
The models will keep changing every few months. A good evaluation and testing discipline and system, built for your task, keeps everyone honest.
You don't win at AI with the cleverest prompts. You win by knowing what good enough means for your work, and measuring it.
#AI #AIRiskManagement #Evaluation #VentureCapital

