Databricks shipped a SQL function that makes structured decisions on unstructured data in sub-second latency. It returns probabilities, not paragraphs — and for a surprisingly large class of enterprise problems, that’s exactly the right tradeoff.
The LLM hammer problem
Here’s a pattern that has become depressingly common in 2026: a team needs to classify customer reviews into five categories. They spin up a managed LLM endpoint. They write a prompt that says “classify this review into one of the following categories.” They parse the response string. They add retry logic for when the model returns something outside the expected set. They add latency budgets because the LLM takes 2-4 seconds per call. They add cost monitoring because they’re burning tokens on a task that doesn’t require a single word of generated text.
The classification itself takes 50 milliseconds of actual decision-making. The other 3,950 milliseconds are overhead from using a text generation model for a task that has nothing to do with text generation.
This is the LLM hammer problem. When your only tool generates language, every structured decision looks like a language generation task. And the cost compounds: latency per call, tokens per call, parsing complexity per call, multiplied across millions of rows per month. Enterprise teams are spending real money — and real engineering time — wrapping text generation around problems that were never about text.
| What the task actually needs | What the LLM provides | Wasted overhead |
|---|---|---|
| A category label | A paragraph explaining the category | Token generation, parsing |
| A probability score | A confidence statement in English | String-to-number conversion, calibration guesswork |
| A ranked score on a scale | ”I would rate this a 4 out of 5 because…” | Output validation, retry logic |
ai_decide is Databricks’ answer to this specific problem. Not a general-purpose LLM function. Not another wrapper around text generation. A purpose-built SQL function that takes unstructured text, evaluates it against a set of questions, and returns structured decisions directly — probabilities, categorical choices, or scores on ordered scales. The output is already structured. There’s nothing to parse.
What ai_decide actually is
ai_decide launched in beta on September 30, 2026 as a native Databricks AI function. It’s callable in SQL and via REST API. Under the hood, it’s powered by TypeSafe AI’s Jev model — a decision model, not a language model.
A language model predicts the next token in a sequence. It generates text. You can coerce it into returning structured output, but that’s a workaround, not its design purpose. A decision model takes unstructured input and a set of questions and returns decisions and their probabilities directly. Structured output is the native format — not a constrained subset of a larger capability.
For each question you pose to ai_decide, you get back one of three output types:
| Output type | What it returns | Use case |
|---|---|---|
| Probability | Numeric likelihood (0-1) | “How likely is this review describing a defective product?” |
| Choice | One of your named categories | ”Is this ticket: billing, technical, account, or other?” |
| Score | Position on an ordered scale | ”Rate the urgency of this support request: low, medium, high, critical” |
These aren’t generated strings that happen to contain numbers. They’re native structured outputs from a model that was designed to produce them. The difference shows up in three places that matter at enterprise scale: latency, cost, and reliability.
The SQL interface
The function lives in SQL. You call it like any other function in a query — same place your transformations already run, against the same governed data, with no separate inference endpoint to stand up.
SELECT
review_id,
review_text,
ai_decide(
review_text,
ARRAY(
NAMED_STRUCT(
'question', 'What is the primary issue category?',
'type', 'choice',
'options', ARRAY('product_quality', 'shipping', 'pricing', 'customer_service', 'other')
),
NAMED_STRUCT(
'question', 'Is the customer likely to return the product?',
'type', 'probability'
),
NAMED_STRUCT(
'question', 'How urgent is the customer sentiment?',
'type', 'score',
'options', ARRAY('low', 'medium', 'high', 'critical')
)
)
) AS decisions
FROM product_reviews
WHERE review_date >= '2026-09-01';
One function call. Three structured decisions per row. No prompt engineering. No output parsing. No retry logic for when the model decides to be creative with your category labels.
Compare that to the LLM equivalent:
-- The LLM approach: generate text, hope it's parseable
SELECT
review_id,
ai_query(
'databricks-meta-llama-3-3-70b-instruct',
CONCAT(
'Classify this review into exactly one of these categories: ',
'product_quality, shipping, pricing, customer_service, other. ',
'Also estimate the probability (0-1) that the customer will return the product. ',
'Also rate urgency as: low, medium, high, or critical. ',
'Return ONLY a JSON object with keys: category, return_probability, urgency. ',
'Review: ', review_text
)
) AS raw_llm_output
-- Now parse the JSON string... handle malformed output... retry on failure...
FROM product_reviews
WHERE review_date >= '2026-09-01';
The first query returns structured data. The second returns a string that you hope contains structured data. At 10,000 rows, the difference is annoying. At 10 million rows, it’s architectural.
Where this fits in the Databricks AI function stack
Databricks has been building out SQL-native AI functions for the past year. ai_decide doesn’t replace the existing functions — it fills a gap between them and general-purpose LLM inference.
| Function | What it does | Output type | Best for |
|---|---|---|---|
ai_classify | Assigns a label from a provided set | Single string label | Simple, flat classification |
ai_analyze_sentiment | Returns sentiment polarity | Positive / negative / neutral | Sentiment-only workflows |
ai_extract | Pulls named entities from text | Structured entity set | NER extraction |
ai_query | General-purpose LLM inference | Free-form text | Summarization, translation, open-ended generation |
ai_decide | Multi-question structured decisions | Probabilities, choices, scores | Compound classification, routing, scoring, evaluation |
ai_classify takes text and returns one label from a set. It’s a single classification. ai_decide takes text and returns answers to multiple questions simultaneously — and those answers can be probabilities, choices, or scores, not just labels. It’s a multi-dimensional decision engine, not a single-axis classifier.
If you need “is this review positive or negative?” — use ai_analyze_sentiment. If you need “categorize this ticket” — use ai_classify. If you need “categorize this ticket, estimate resolution time, assess customer churn risk, and score the priority” — that’s ai_decide. Four questions, one function call, four structured outputs.
The real-time API: agents that decide, not generate
The SQL interface handles batch workloads. But ai_decide also ships with a REST API for real-time applications — and for teams building agentic systems, that changes the routing layer entirely.
Consider a common agent pattern: an orchestrator receives a user request and needs to decide which specialist agent to route it to. The classic approach uses an LLM to “reason” about the routing — which means generating text, parsing it, and hoping the model picks the right route. With ai_decide, the routing decision is a sub-second structured call:
import requests
def route_request(user_message: str) -> dict:
response = requests.post(
"https://<workspace>.databricks.com/api/2.0/ai-decide",
headers={"Authorization": f"Bearer {token}"},
json={
"input": user_message,
"questions": [
{
"question": "What is the reasoning complexity?",
"type": "choice",
"options": ["simple_lookup", "multi_step_reasoning", "creative_generation"]
},
{
"question": "Which domain does this belong to?",
"type": "choice",
"options": ["billing", "technical", "account_management", "product"]
},
{
"question": "How likely is this to require human escalation?",
"type": "probability"
}
]
}
)
return response.json()
The routing decision comes back in milliseconds with structured data — not a paragraph that you parse into a routing decision. For agent systems processing thousands of requests per hour, the latency and cost difference compounds fast.
This maps directly to the model-routing pattern described in the Databricks announcement: a software company’s AI assistant uses ai_decide to assess each incoming prompt’s reasoning difficulty, then routes simple lookups to a cheap, fast model and complex reasoning tasks to a more capable (and expensive) one. The router itself doesn’t need to reason. It needs to decide. Different problem, different model.
AI-as-judge without the judge’s overhead
One of the most expensive uses of LLMs in production is evaluation — using an LLM to judge the quality of another LLM’s output. Every major agent framework does this: generate a response, then call a second LLM to evaluate whether the response was correct, complete, and policy-compliant. The evaluator LLM is often the same size as the generator, doubling the cost and latency of every interaction.
ai_decide flips this. Instead of asking a second LLM to write an evaluation essay that you parse into a score, you ask a decision model to score directly:
SELECT
response_id,
generated_answer,
reference_policy,
ai_decide(
CONCAT('Answer: ', generated_answer, ' | Policy: ', reference_policy),
ARRAY(
NAMED_STRUCT(
'question', 'Does the answer comply with the reference policy?',
'type', 'choice',
'options', ARRAY('compliant', 'partial', 'non_compliant')
),
NAMED_STRUCT(
'question', 'How complete is the answer relative to what the policy requires?',
'type', 'score',
'options', ARRAY('missing_critical_info', 'partial', 'adequate', 'comprehensive')
),
NAMED_STRUCT(
'question', 'Does the answer contain information not supported by the policy?',
'type', 'probability'
)
)
) AS eval_results
FROM agent_responses
WHERE eval_status = 'pending';
Three evaluation dimensions per response. The compliance check, the completeness score, and the hallucination probability all come back as structured data. You don’t parse a paragraph into a verdict. The verdict is the output.
For teams running evals at scale — thousands of agent responses per day — the cost difference between an LLM judge and a decision model judge is significant. The latency difference is even more significant for real-time evaluation loops where the eval result determines whether to show the response or retry.
The open-weight angle
Databricks isn’t locking ai_decide to a single model. The announcement references “a broad and growing ecosystem of open-weight decision models” that can be served and run directly in SQL on Databricks. TypeSafe AI’s Jev model is the default, but the architecture supports alternatives.
The obvious win is data residency. Enterprise teams that can’t send data to external model endpoints — whether for compliance reasons or competitive sensitivity — can serve decision models entirely inside their own Databricks environment. The data never leaves the workspace. But the longer-term play is more interesting: decision models are a category, not a product. Databricks is positioning ai_decide as the SQL interface to that entire category. As the ecosystem of purpose-built decision models grows (and it will — classification, routing, and scoring have never needed text generation), the function surface stays the same while the models underneath get better and more specialized.
Where it falls short
Beta status means beta caveats. ai_decide is not GA. There are no published SLAs, no guaranteed uptime commitments, and no production support tier. Teams building mission-critical pipelines on it today are accepting risk.
No published benchmarks. The announcement says “sub-second latency” and “lower cost than an LLM” but provides zero concrete numbers. No latency percentiles. No cost-per-decision figures. No accuracy benchmarks against LLM-based classification. For a function aimed at enterprise-scale workloads, the absence of hard numbers is notable. Data engineers making build-vs-buy decisions need actual figures, not qualitative claims.
Limited documentation on edge cases. What happens when the input text is ambiguous? How does the model handle multilingual content? What’s the maximum input length? How deterministic are the outputs across identical inputs? The current documentation doesn’t address these questions in depth. For batch pipelines processing millions of rows, edge-case behavior at the 99th percentile matters.
No fine-tuning story yet. The open-weight model ecosystem is mentioned but not detailed. Can you fine-tune a decision model on your domain-specific data? Can you bring your own model that matches the ai_decide interface? The documentation links to an “open-Jev walkthrough” but the customization path isn’t clear.
The category design is yours to get wrong. ai_decide returns structured output, but the quality of that output depends entirely on how well you define your questions and categories. A poorly defined choice set produces confidently wrong classifications just like a poorly written prompt does. The failure mode is different — structured garbage instead of unstructured garbage — but the root cause is the same: bad task definition.
When to use ai_decide vs. ai_classify vs. ai_query
The decision tree is simpler than it looks:
Use ai_classify when you need a single label from a flat list and nothing else. One question, one answer, one category. It’s simpler to call and sufficient for single-axis classification.
Use ai_decide when you need multiple structured decisions per row — compound classification, probability estimation, ordinal scoring, or any combination. Also use it when you need the output as native structured data, not parsed text. Also use it for real-time routing and evaluation where latency matters.
Use ai_query when you need generated text — summaries, translations, explanations, or any task where the output is language, not a decision. ai_query is the general-purpose LLM tool. ai_decide is the purpose-built decision tool. Using ai_query for decisions is like using a word processor to do arithmetic — it can, technically, but the tool doesn’t match the task.
Need generated text?
└── Yes → ai_query
└── No → Need a single label?
└── Yes → ai_classify
└── No → Need multiple decisions / probabilities / scores?
└── Yes → ai_decide
What this signals about the AI function roadmap
ai_decide is the clearest signal yet that the Databricks AI function stack is moving toward task-specific models, not general-purpose LLMs for everything. The trajectory:
ai_analyze_sentiment— a task-specific function for a specific NLP taskai_classify— a slightly more general task-specific functionai_extract— entity extraction as a native SQL operationai_decide— multi-dimensional decision-making with a purpose-built model
The pattern is decomposition. Instead of routing every AI task through a general-purpose LLM and parsing the output, each task type gets a native function backed by a model optimized for that specific task. The SQL interface stays uniform. The underlying model architecture changes to match the workload.
Database engines figured this out decades ago: specialized execution paths for different query types outperform a single general-purpose executor. The AI function stack is learning the same lesson.
Here’s the thing nobody’s saying out loud: most “AI-powered” data pipelines in production right now are using a $0.015/1K-token text generation model to do work that a $0.001/call decision model could handle better. The waste isn’t theoretical. It’s in your cloud bill every month.
For data engineers, the SQL-native AI function surface area will keep growing. The engineer who reaches for ai_query by default for every AI task is stuck in 2025. The one who picks ai_decide for decisions, ai_classify for labels, and ai_query only when they actually need generated text is building pipelines that cost less and break less.
Try it
ai_decide is available in beta on Databricks workspaces. The SQL function is documented in the Databricks SQL reference. The REST API is documented at the standard Databricks API endpoint.
For teams already using TypeSafe AI’s API, the ai_decide function is directly compatible — the same model, native in SQL.
- Databricks ai_decide SQL documentation
- Databricks announcement blog post
- TypeSafe AI / Jev decision model
Related articles
- Your Data Stack Wasn’t Built for This — how AI functions like
ai_decidefit into the broader AI-native data architecture - Databricks Agent Bricks Is Quietly Changing How Data Engineers Work — the agent layer that
ai_decidecan power routing and evaluation for - Snowflake and Databricks Converged on AI Agents in 2026 — how both platforms are building the same agent stack
If you want to understand the foundational patterns behind data-intensive systems — how storage engines, encoding formats, and distributed architectures work under the hood — Designing Data-Intensive Applications by Martin Kleppmann covers the principles that make purpose-built execution models like decision engines possible in the first place.
Disclaimer: This article is based on Databricks’ public announcement and documentation for ai_decide as of September 30, 2026. The function is in beta and subject to change. The author has no affiliation with Databricks or TypeSafe AI. No independent benchmarks were available at the time of writing — all performance claims (sub-second latency, lower cost than LLMs) are from Databricks’ announcement and have not been independently verified. This article contains affiliate links — purchasing through them supports this blog at no extra cost to you.


Comments
Loading comments...