10 growth jobs I'd hand to a model that only makes decisions
I tested Jev on two small pieces of growth work this weekend. It went well, and it barely scratched the surface. Here are the ten builds I think matter most.

A new AI model called Jev launched last week and the internet decided it does everything. One well-known developer posted that it had replaced ChatGPT, Cursor, Spotify and McDonald's for him. He was joking. A fair few of the people reposting him were not.
I build AI growth systems for founders, so I had one question: does it make growth work better, or is it just fast?
If you watched my Reel, this is the longer version with the parts I couldn't fit into three minutes. We'll cover what Jev is, what happened when I tested it, why those tests were the small version of the idea, and the ten places I think it could change how growth runs. Then the caveats, because there are a few.
If you're new here, 5 steps to having AI run growth systems for you covers the order I'd build in, and What is growth engineering? covers what you're building. This one is about a new part you can put underneath both.
First, what Jev actually is
Jev is made by TypeSafe, founded by Diogo Almeida, who worked at OpenAI on the methods that made language models follow instructions. They call it a System One model, which is their own name for the category. It writes nothing, chats with nobody and browses nothing. You give it a situation and a question, and it gives you a decision: yes or no, one option from a list you wrote, or a score. Every answer comes with a probability.
TypeSafe's own description is that language models generate, and Jev judges. Their launch post prices input at $0.042 per million tokens, against $0.20 to $10 for frontier models. It quotes end-to-end responses of 70 to 500 milliseconds, against 3 to 329 seconds, which they put at 40 to 200 times faster. In my tests it averaged about 0.4 seconds. Those are the vendor's numbers. Mine are below.
There's no chat box. You reach it through code, every question needs a fixed answer type, and you have to write your own definition of yes and no. That last part is where most of the work is, and I'll come back to it.

A language model does the writing. Jev does the deciding. They're different jobs, and until now we've been paying the writer to do both.
What I tested
I built a small site that holds 998 real startups from Y Combinator's public directory. To be clear about who did what: Jev didn't go and find these companies or research them. I gave it the dataset, and it judged every company against the same criteria. It only sees what each company says about itself. It never sees the name, so reputation can't help.
Test one: who needs growth help right now? I gave it three yes or no questions with my own definitions: is winning customers still the founders' own job, do they sell to businesses today, and could they pay about €10,000 for outside help.
It judged all 998 in 22.4 seconds for 3.1 cents, multiplied the three probabilities and handed me a top 20. Mostly two and three person teams selling to hospitals, trading firms and chip designers. That is almost exactly who buys from me, which was slightly unnerving.
It also marked 376 companies as unsure and left them for a person to read. I like that more than I expected to. It only knows what the directory tells it, and it says so.
I priced the same job through a frontier model from a five company sample. It came to roughly $10, so about 300 times the cost on this one job. On speed the gap was smaller, around six times per call. That's one measurement on one task, and I wouldn't quote it as a universal number.
Test two: how easy it is to get this wrong. I typed my own question: does this company have a marketing team bigger than two people? It ran all 998 in 40 seconds and gave me a confident looking list. Nothing on it scored above 0.69. The directory doesn't list marketing teams, so it was inferring from headcount.
That's the lesson I'd put on a wall. Jev scales whatever thinking you hand it. A lazy question gets you the same wrong answer, faster and cheaper, across your whole market.
Both tests were batch jobs. I pointed it at a list and pressed a button. Useful, and also the smallest version of what this is for.
The shift: stop asking once a day
The big unlock is making judgment cheap enough that you can put it inside every part of a growth system. We already had plenty of AI that creates things.
Here's what changes when a decision costs a fraction of a cent and takes under half a second. You stop asking one big question each morning ("what should we do today?") and start asking hundreds of small ones all day, inside the system, without a person typing anything.
The pattern I'd build around looks like this:

Code calculates the facts. Jev makes the judgment calls. Agents do the work. Jev checks the work. A person sees the few things that need them.
TypeSafe make the same split in their own examples: arithmetic, dates and statuses stay in code, and the model only gets the judgment. I'd follow that. Asking any model to do your sums is how you end up with confident nonsense.
I haven't built all ten of the ideas below. I've tested pieces of two. I'm sharing the list because I think this is where growth systems go next, and because it's more useful to argue about a specific list than a vague one.
1. A supervisor for your AI's work
If you run AI agents, someone has to check what they produce. Today that someone is you.
Before any draft reaches you, Jev can answer: is every claim in this supported by the source? Is this the right person at the right company? Does it match the conversation so far? Is there anything in here a stranger planted? I tested that last one by accident. A prospect's profile contained a hidden instruction written for AI assistants, and Jev flagged it at 96 percent.
Confident passes go to your approval queue. Anything unsure gets read by a person first. Nothing sends itself.
What Jev decides: pass, fix or escalate. What you still do: approve what goes out.
2. Account state instead of a lead score
"Lead score: 73" tells you nothing you can act on. I'd rather know that an account is problem-aware, evaluating technically, and stalled on an integration concern.
Every new event (a page visit, a reply, a call note) triggers a handful of small questions: does this still fit, what stage are they at, what's the likeliest objection, who should act next, and are we confident enough to act at all.
What Jev decides: the state and the next owner. What you still do: have the conversation.
3. What should this person see next?
A visitor from LinkedIn, 18 person fintech, third visit, read the pricing page and a case study. What should your site show them? A demo button, a calculator, another case study, or nothing?
This is journey orchestration, and it needs an answer before the page loads. That rules out a slow model and suits a fast one. It's also a bounded choice from a list you wrote, which is exactly the shape Jev handles.
What Jev decides: intent, interest and the next best module. What you still do: write the modules.
4. Customer signals that turn into work
Calls, support tickets, lost deal notes, reviews, replies. Each one gets labelled: pain point, objection, positioning signal, product friction, competitor mention, content idea, retention risk.
Then code counts them. When seventeen separate people struggle to understand your pricing, you have a pricing page test, an explainer and a sales note, each routed to whoever owns it.
What Jev decides: what each signal is and where it goes. What you still do: decide what's worth fixing.
5. Experiments that know when to stop
Your code works out the numbers: click rate, conversion, cost, sample size, significance. Jev handles the fuzzy calls a growth person makes while staring at those numbers. Continue, scale, pause, kill or segment? Is this creative fatigue or a weak offer? Is there enough evidence to change direction? Does a person need to look?
The writing model only gets involved after that decision, to produce the next test.
What Jev decides: what happens to each test. What you still do: set the guardrails and the budget.
6. Creative that evolves instead of multiplying
"Generate 50 ads" gives you 50 ads and no idea why any of them worked.
Treat each ad as a set of parts: audience, hook, pain, promise, proof, format, ask. Jev judges which part failed and which should change next. The brief to the writing model becomes specific: keep the audience and the proof, replace the hook, move from fear to aspiration, give me five.
What Jev decides: what to keep and what to mutate. What you still do: approve what runs.
7. Lifecycle that decides per person
Day 1 email, day 3 email, day 7 email is a schedule. Nobody checked whether this particular person needs an email today.
For every user, regularly: do nothing, educate, activate, remind, rescue or hand to a human? Through which channel? With what objective? Then a model writes the message. "Do nothing" being a real option is most of the value here.
What Jev decides: the next best action per person. What you still do: own the rules on how often anyone hears from you.
8. A backlog that writes itself
Feed it the day's observations from analytics, CRM and ads. Each one gets judged: unusual or normal, actionable or not, known issue or new, how strong is the evidence, which growth lever does it touch.
Add them up and the backlog arrives with evidence attached: "referred users activate at twice the rate, and almost nobody is asked to refer". That beats whoever spoke last in the meeting.
What Jev decides: what deserves attention. What you still do: choose what to build.
9. A growth control room
This is the one I'd most like to build. One system watching product events, CRM, ads, search, email and calls, asking small questions continuously. Is acquisition getting worse or is this normal variance? Which stage is responsible? Which playbook applies? Can an agent handle it, or does a person need to see it?
Your morning starts with a short summary: three things noticed, two handled, one experiment started, one decision waiting for you.
What Jev decides: what's happening and who owns it. What you still do: the one decision that's waiting.
10. Growth reflexes
We talk about AI agents as things you ask. The more interesting idea is a business that reacts without being asked.
Traffic quality drops and acquisition adjusts. The same objection shows up three times and a piece of content gets drafted. A high intent account appears and the right person hears about it within the minute. A winning ad emerges and budget moves toward it.
I think of it as a growth nervous system: notice, decide, act. The model doing the deciding is the part that's been missing, because until now every decision meant a slow, expensive call to a model built for writing.
What Jev decides: hundreds of small things. What you still do: set what it's allowed to act on alone.

The caveats, because there are a few
It can be wrong. TypeSafe say it can't hallucinate. What that means is it can't wander off and invent an answer, because you've fixed the possible outputs in advance. It can still pick the wrong one, and their own FAQ says so. My marketing team test is the proof. The probability is how you catch it: set a band, and anything inside it goes to a person.

The question is the work. Writing a precise definition of yes and no, with examples, is harder than it sounds. It's the same job as teaching your judgment in step two of the 5 steps guide. Jev just makes it unavoidable.
It only knows what you show it. If the data can't answer the question, you get a guess with a number on it.
It needs a builder. No chat box means no trying it over lunch. You need someone comfortable with an API.
Mind where the data goes. Every call sends something to a third party. Company facts are low risk. A customer's own words are not. Send the minimum the question needs.
What I'd do first this week
Pick one decision you make fifty times a week. Is this lead worth a reply? Is this comment a buyer? Does this draft make a claim we can't support?
Write your definition of yes and no. Add three examples of each. Then collect a hundred past cases where you already know the right answer.

Run them through and compare its answers with yours. Where it disagrees, one of you is wrong, and it's more often the definition than the model. Fix the definition and run it again.
You'll learn more from that than from wiring up a control room on day one.
This is the work I do with founders at FounderGrowthOS: working out which decisions should come off your plate, building the system around your judgment, and keeping it honest as it grows.
If you'd like help with yours, reply with the decision you make most often and how you make it today. That tells me more than the name of the tool does.
Written by Amy Wilkinson, 21 September 2026.
Want a second opinion on the system you're building?
A 20-minute call and a free growth diagnostic you keep either way, to run yourselves or with us.
Book the call →