Across 915 answers to 30 buying questions, a steep power law emerges. MLflow and Kubeflow win both memory and retrieval. The challengers win only one door — if any.
An industry study by Lil Big Things.
Share of AI answers — the percentage of all 30 category prompts where each brand was named, across 6 engines and 915 answers.
| # | Brand | Share | ChatGPT | Perplexity | Google AI | Gemini | Claude | Copilot |
|---|---|---|---|---|---|---|---|---|
| 1 | MLflow | 42.1% | Strong | Strong | Strong | Strong | Strong | Strong |
| 2 | Kubeflow | 32.6% | Strong | Moderate | Moderate | Strong | Strong | Moderate |
| 3 | Weights & Biases | 20.2% | Moderate | Weak | Moderate | Moderate | Strong | Moderate |
| 4 | ClearML | 12% | Moderate | Weak | Moderate | Moderate | Weak | Moderate |
| 5 | Comet ML | 2.8% | Absent | Weak | Absent | Weak | Moderate | Weak |
| 6 | ZenML | 2.8% | Weak | Absent | Weak | Weak | Moderate | Weak |
Want this study as a designed PDF?
Enter your email for the full PDF and the raw dataset. The study stays free to read either way.
You will also get each new study. Unsubscribe anytime.
Shipping a model to production touches a dozen concerns: tracking experiments, versioning data, scheduling GPUs, orchestrating pipelines, and serving the result. The vendors solving this sell overlapping slices under different banners — experiment tracking, pipeline orchestration, GPU scheduling, model registry, deployment. Teams rarely describe the need the same way twice, and most stitch together two or three tools to cover it.
That sprawl is the defining trait of this category, and it is why AI assistants behave differently here than a keyword search does. Ask a search engine for "experiment tracking tool" and you get a tidy list of trackers. Ask an assistant "what open-source platform handles tracking, GPUs, and deployment?" and you get the most-written-about names in open-source ML, weighted heavily toward the two or three every tutorial mentions.
Two competitive sets exist. The search set — who ranks on Google for the category keywords — is broad. The answer set — who AI names when buyers describe the problem in their own words — is narrow and top-heavy. This report measures the answer set, because that is the room where the category is now being decided.
Every prompt was split by buying stage. This category skews heavily to exploration: 28 of the 30 tracked questions are awareness-stage, with only two pushing toward a decision. That itself is a finding.
At the awareness stage — "what open-source platforms track experiments?", "what tools manage GPU clusters?" — the field is wide and generous. Assistants list five or six names, and ClearML surfaces reliably. The two consideration questions in the set, both about managing and cost-optimizing GPU clusters for specific client projects, narrow the field toward whichever platforms have concrete, comparative write-ups.
Because demand is front-loaded into exploration, awareness presence is the prize. The names that get listed when a buyer first asks "what open-source options exist?" shape the shortlist before any comparison happens. In this category, being in the first list an assistant produces matters more than winning a later head-to-head.
There is no single AI answer. Presence varies dramatically by assistant, and the variation is not noise — it tracks how each engine builds its answer.
Sort the engines by how they answer and the split is stark. Engines that search the live web in real time — Google AI Overviews, Gemini, and Copilot — reward fresh documentation and third-party write-ups. Engines that lean on trained memory or tight citation sets — Claude and Perplexity — reward an established footprint.
The same brand can score very differently across the two types. In this category, ClearML reaches 26% on Google AI Overviews and 1.7% on Perplexity, while Weights & Biases reaches 52% on Claude and 8% on Perplexity — a six-fold gap. Where a brand is strong tells you how it earned its presence.
A brand's fitness for the job barely moves the memory engines — they name who they were trained to know, and the open-source ML canon is dominated by MLflow and Kubeflow. But the retrieval engines can be influenced in weeks with the right content and citations. That asymmetry — a fast lane and a slow lane — is the strategic backdrop for every brand in the category.
| Engine | Type | ClearML share | Characteristic |
|---|---|---|---|
| Google AI Overviews | Retrieval | 26.2% | ClearML strongest surface — searches live web |
| Gemini | Retrieval | 24.6% | Strong on tracking and orchestration prompts |
| Copilot | Retrieval | 19.2% | Solid across integrated-tooling questions |
| ChatGPT | Memory-led | 11.7% | Most-used assistant, mid-pack for ClearML |
| Claude | Memory only | 4.2% | Names incumbents from training — ClearML thin |
| Perplexity | Citation-led | 1.7% | ClearML nearly absent from cited sources |
ClearML per-engine presence rates, illustrating the retrieval vs. memory engine split across the MLOps category.
Every appearance maps to a real question. Read across the six question themes and a clear pattern emerges: MLflow and Kubeflow are near-ubiquitous, the focused trackers concentrate in tracking and open-source themes, and the broad platforms spread their marks across more of the workflow.
MLflow shows consistent presence in tracking, open-source, and workflow themes, with partial coverage in GPU orchestration, deployment, and cost. Kubeflow is strongest in GPU orchestration, deployment, workflow, and cost — the infrastructure end. Weights & Biases concentrates in tracking and open-source themes and is absent from GPU and cost questions.
ClearML has the broadest theme coverage: consistent presence in tracking, GPU orchestration, workflow, and open-source, with partial coverage elsewhere. This is the integrated platform profile — it shows up when the question asks for several things at once. That is its best ground, and the clearest opportunity for a challenger to win contested prompts.
The content thesis for the category: specificity beats fame on narrow queries. The brands that win contested prompts are the ones whose documentation is unusually exact about a capability — and whose exactness is echoed on third-party pages. Generic breadth loses these; precise depth wins them.
When an engine searches to answer a category question, approximately 96% of cited URLs are third-party — documentation, tutorials, blog round-ups, GitHub threads, and forums. About 3% point to competitor-owned domains. Under 1% point to any single vendor own domain: no brand site is a meaningful source of its own AI citations.
AI assistants do not answer category questions from a vendor website. They answer from the consensus of the open web — documentation, tutorials, GitHub threads, and comparison posts — and then, if a brand is famous enough, from memory. This is why MLflow and Kubeflow win: years of tutorials, Stack Overflow answers, and "top MLOps tools" listicles about them are the evidence base. A platform own docs, however good, are weighted far below independent corroboration.
On Perplexity specifically, third-party citation share exceeded 96%. A live Claude answer to an open-source tracking question named six tools and cited none — answered entirely from training memory, which is why memory-resident names appear and retrieval-led ones do not.
The lesson generalizes to every brand in the category: to be named in an answer like this, you do not need a better page about yourself. You need to be named inside the third-party pages the engines learn from — the tutorials, round-ups, and repositories that train what the model recalls.
Reading the raw responses side by side exposes the real structure of AI visibility in this category. An assistant names a brand through memory or through retrieval — and the two doors need entirely different keys.
Door 1 is memory. Opened by fame, over years. Engines answering from training data name tools they know. They reach for the names with a massive footprint of tutorials and repositories — MLflow, DVC, Weights & Biases, Metaflow. ClearML is not yet dense enough in that corpus to be recalled by default. A live Claude answer named six tools and cited none: MLflow, DVC, Weights & Biases, Kedro, Aim, Metaflow. Zero citations. Pure recall. ClearML absent.
Door 2 is retrieval. Opened by content, in weeks. Engines that search live name tools they find in top sources. A well-documented platform surfaces well here even without deep fame — as long as its docs and third-party write-ups get retrieved. This is where the broader mid-field platforms do their best. Google AI Overviews and Gemini give mid-field platforms 20 to 26% while they score far lower on the memory engines.
MLflow and Kubeflow win both doors: they are famous (memory) and heavily cited by third parties (retrieval). Weights & Biases wins the memory door far more than the retrieval one. The mid-field platforms win retrieval and lose memory. Each brand position on these two doors is the single best predictor of where its presence comes from and what it must do next.
The six platforms sort into clear positions based on how they win the two-door question. Both-door brands (MLflow, Kubeflow) are strong everywhere. Single-door brands split into two groups: memory-led (Weights & Biases, Comet ML, ZenML) and retrieval-led (ClearML).
The both-door brands are effectively uncontestable on their home turf. MLflow peaks at 65% on Claude — the strongest memory signal in the study — meaning it is deeply baked into what the models recall as well as what they retrieve. For any challenger, competing with MLflow on tracking alone is the losing move.
The memory outlier is Weights & Biases: 52% on Claude but just 8% on Perplexity, a six-fold gap. Its presence is reputation-driven — carried by how much has been written about it over time, rather than by fresh, citation-friendly content. Strong where models recall; thin where they cite live.
ClearML is the mirror image: 26% on Google AI Overviews and 25% on Gemini but only 4% on Claude and 2% on Perplexity. Its breadth wins integrated, multi-capability questions when engines search live, but it has not yet accumulated the memory footprint the leaders hold. Which door a brand wins tells you whether its presence comes from accumulated reputation or from current content — and therefore what it must do to grow.
An entity-rich profile of each brand studied — how they are positioned, where they win, and where they do not.
What brands in this category should do differently, based on what the data shows.
Seed the third-party record. Get named in the "top open-source MLOps tools" round-ups, comparison posts, framework docs, and GitHub-adjacent tutorials that engines cite and learn from. This attacks the 96% third-party citation share directly and is the single highest-leverage move for both doors.
Match content to the door you are losing. Retrieval-led brands need sustained presence in cited sources to train the memory engines. Memory-led brands need fresh, structured, citation-friendly content to win live retrieval. Diagnose which door is shut before spending.
Own the integrated question. Broad, end-to-end platforms beat narrow ones when a buyer asks for several capabilities at once. Publish and earn coverage on multi-capability queries rather than fighting the category leaders on a single capability they already own.
Press the retrieval engines for fast wins. Google AI Overviews, Gemini, and Copilot respond to content and citations in weeks. They are the fast lane for any brand trying to move its numbers this quarter.
Work the memory door patiently. Claude and Perplexity move only as sustained third-party presence trains into future models and tightens citation sets. Treat this as a quarters-long program, not a campaign.
Pick a reachable target, not the leader. The category is a steep power law. For most brands the realistic contest is with the name one tier up, on a clear axis of difference — not a head-on fight with MLflow or Kubeflow on their home turf.
This study ran 30 non-branded seed questions across 6 AI engines between 5 – 12 Jul 2026, capturing 915 machine-generated answers. Category structure and the organic competitive set were derived from Ahrefs. AI-visibility metrics were measured in Scrunch AI.
Prompts are non-branded, written the way a real buyer describes the problem in their own words. We distinguish between awareness prompts and evaluation prompts. Figures are directional estimates from third-party measurement, not audited results.
Want this run for your category?
We can measure your market the same way and show you exactly where you stand in AI answers.
Answer Engine Index by Lil Big Things · July 2026 · All study content is open and crawlable. Methodology