EP 181: Jim Wrubel – Surfacing the Right Info at the Right Time
- Jim and Tim haven’t met yet, but Jim was tagged on a LinkedIn post by someone that Tim respects.
- Jim comes from the Marketing world and has a product that helps companies get discovered by LLM tools like ChatGPT, Claude, Gemini, etc.
- What does Marketing have to do with ERP? It’s all about getting the right information discovered at the right time, whether you’re going out to the public market to look for something or whether you’re looking for something in your internal ERP.
AI Summary (written by AI, edited by Tim)
AI Search Is Still Search: What That Means for Marketing, ERP, and Private Data
The most useful thing Tim Rodman learned from Jim Ruble was that AI search is a lot less magical than it looks. ChatGPT can feel like an oracle sitting on top of the entire Internet, somehow knowing which companies, products, articles, and ideas are worth recommending. But Ruble’s description was much more mechanical: understand the user’s intent, turn that intent into searches, retrieve candidate information, rank it, read the best sources, and assemble an answer.
That distinction matters because it connects two problems that initially look completely different. One is public: how does a company get discovered and recommended by ChatGPT, Claude, Copilot, or another AI assistant? The other is private: how does someone inside a company ask a natural-language question and get the right answer from ERP data, SharePoint, Google Drive, a BI system, policies, procedures, or an internal knowledge base?
In both cases, the hard problem is still retrieval. The LLM is very good at understanding what someone is trying to accomplish. But it still needs somewhere to go for the underlying information. That was the thread running through the entire conversation.
And for Tim, it changed how he was thinking about AI in a pretty fundamental way.
AI Search Is a Thin Layer on Top of Search
Ruble likes to describe the “dirty secret” of AI as a thin layer on top of search. If someone asks ChatGPT a question that needs current information, the model does not necessarily reach into some giant internal brain and pull out the answer. It may perform one or more web searches, choose results, fetch pages, extract useful pieces, and summarize them back to the user.
That immediately makes AI visibility feel less alien. Tim had gone into the conversation assuming that Answer Engine Optimization, AI Engine Optimization, or whatever acronym eventually wins was mostly a brand-new discipline. Ruble’s estimate was that roughly 70% of the traditional SEO foundation is still relevant, while acknowledging that people can argue about the exact percentage.
Even if that number were only 40% or 50%, Tim’s reaction was basically the same. That is still a lot more overlap than zero.
The consumer has changed.
Instead of optimizing only for a human who searches Google, clicks a blue link, lands on a page, and reads it, companies now also have to think about a bot that searches, fetches the page, extracts the relevant passage, and may never send the human visitor to the website at all.
One Prompt Can Turn Into Many Searches
The interesting part is what happens between the user’s prompt and the search engine. A person might type something conversational such as asking for wireless headphones for a daughter. The AI can turn that into multiple narrower searches around reviews, budget options, over-the-ear designs, noise cancellation, and other characteristics implied by the request.
Those are commonly called fan-out queries.
The number of fan-out queries can vary. Ruble said a lighter or free interaction might produce only one or two, while deeper reasoning modes or higher-tier models could produce 20 to 40. The exact behavior depends on the model and situation, which is part of why trying to reduce AI search to one predictable keyword is difficult.
The AI then has to decide which search results deserve more attention. Ruble described a re-ranking process where a larger set of candidates might be narrowed to a handful of pages worth fetching in full. The model can then pull chunks from those pages, determine which combination best answers the question, and build the response.
That happens fast.
Tim found that surprisingly reassuring. Not because the technology is simple, but because the mental model is understandable. The model is not climbing a mystical mountain to consult the Oracle of Bluetooth Headphones. It is doing retrieval, ranking, selection, and synthesis at high speed.
Once that clicks, a lot of the surrounding behavior makes more sense.
Why Google Analytics Can Miss AI Traffic Completely
Ruble’s path into AI visibility started with a lead that appeared to come out of nowhere. In February 2024, a prospect told him that ChatGPT had recommended his company. Ruble was watching his website analytics closely enough that this made no sense. The prospect had somehow discovered the company, visited the site, and booked a demo without leaving the trail Ruble expected to see.
His immediate response to the prospect was essentially, “You’re lying.”
Not a recommended sales technique.
But the failed interaction sent him down the rabbit hole that eventually became SpyGlasses. He learned that an AI citation fetch does not necessarily behave like a normal human page visit. Google Analytics typically relies on JavaScript running in the visitor’s browser. A bot that just wants the text of the page can fetch that content without executing the JavaScript.
So the AI can read the page while Google Analytics sees nothing.
The underlying web server still sees the request. Ruble went into his raw web logs and looked for the identifying user agent associated with ChatGPT citation traffic. At the time his company was small enough that he could literally scroll through recent logs and find the requests by hand.
He then built tooling around that.
The practical lesson is that a website can have meaningful AI activity that is invisible in normal human-traffic analytics. If the goal is to understand whether AI systems are actually fetching and citing the site, server logs and bot identification matter.
That is a different measurement problem from counting human sessions.
SEO Is Not Dead Because AI Still Needs Something to Find
If AI retrieval starts with search, traditional search visibility still matters. Ruble’s point was simple: if a page never becomes a candidate for the fan-out queries being run, there is effectively no chance for that page to become the cited source.
That does not mean AI optimization is identical to traditional SEO.
It means the foundation did not disappear.
One difference is that the searches generated by an AI can look weird compared with normal human searches. They may be longer, more specific, and constructed around the semantic intent of the original prompt rather than the short phrase a person would normally type into Google.
Tim liked the idea of inspecting those fan-out queries directly. Ruble explained that, at least in some ChatGPT web interactions, the queries can be visible through browser network inspection, and there are Chrome extensions designed to capture and present them in a structured form.
For someone trying to understand AI discovery, that turns an abstract concept into something inspectable.
What did the user ask?
What searches did the model generate?
Where does the company rank for those searches?
Which pages got selected?
That is a much more concrete optimization problem.
The Second Lever Is What Other People Say About You
A company’s own website is only one source the AI can use. Ruble described a second major lever: third-party sources.
That can include review sites, comparison articles, industry blogs, Reddit discussions, and other places where a company or product gets mentioned.
This becomes especially important for recommendations. If someone asks an AI assistant which ERP system would be best for a particular company, the assistant may prefer sources that compare products, discuss experiences, list strengths and weaknesses, or provide some kind of independent signal.
In other words, being great on your own website is not enough.
Tim pointed to ERPSoftwareBlog.com as the kind of third-party site that could matter in the ERP ecosystem. The important part is not that every company needs to be on one particular website. The larger point is that the AI may build its answer from sources the company does not control.
That sounds familiar because it is familiar.
Reviews mattered before AI.
Industry credibility mattered before AI.
Links mattered before AI.
Third-party mentions mattered before AI.
AI changes the way those signals get assembled and delivered, but it did not invent the idea that outside sources influence trust.
Reddit May Matter More Than the Fancy Press Mention
One of the stranger implications involves Reddit. Ruble explained that some large publishers block AI systems from using their content, while Reddit content can be highly accessible and valuable for AI retrieval.
That can flip the old prestige hierarchy on its head.
A full-page feature in a major business publication might feel like the obvious marketing win. But if the AI cannot access or cite that content, it may have little impact on AI discovery. Meanwhile, one ordinary Reddit user saying that a product works well for a particular use case could become evidence in a future AI recommendation.
Tim found that fascinating.
He also finds Reddit kind of miserable to use.
His problem with Reddit was not really the idea of the community. It was the noise, the ads, the interruptions, and the general difficulty of following what was happening. He had tried to get involved several times and bounced off.
Ruble’s advice was not to turn Reddit into another broadcast channel.
Be a Reddit user.
Build karma by contributing useful things.
Disclose who you are and where you work when that matters.
Then answer the question.
No promotion required.
“Good Intent Wins” Is More Than a Nice Phrase
That advice reminded Tim of a Gary Vaynerchuk line he keeps on his wall: “Good intent wins.” In the context of Reddit, the practical version is pretty straightforward. If someone has a real question and you have relevant expertise, be transparent about your position and try to help.
Ruble said useful participation can compound for a long time. His company was still appearing in certain ChatGPT prompts because of a Reddit comment he had made more than a year earlier.
That is an interesting kind of content durability.
The goal is not to sneak a sales pitch into a community. It is to contribute something worth retrieving later. Original data can be especially useful because it gives the answer something more substantial than an opinion. Ruble gave an example of using his company’s own citation data to explain what percentage of sources came from LinkedIn.
If someone finds that helpful, great.
If they want to know who provided the answer, the profile can contain the company information.
For Tim, that was probably the most attractive version of “AI optimization.” Create useful material with good intent, put it somewhere the systems can retrieve it, and make the information strong enough that it deserves to become part of an answer.
There are technical games around the edges.
But the center is still useful content.
You Can Avoid Living Inside Reddit
Ruble also suggested a more tolerable way to participate in Reddit without constantly browsing Reddit. He mentioned a social-listening tool that can monitor keywords and send relevant conversations into Slack, Teams, or another notification channel.
That part immediately clicked for Tim.
The monitoring tool can effectively become the Reddit interface. Instead of scrolling through an endless feed, someone can receive a small excerpt when a relevant discussion appears, click into that specific thread, decide whether they can add value, and leave a useful response.
Then leave.
Much better.
This matters because community participation is its own strategy. Ruble was clear that companies should not expect to casually drop into Reddit once a month, paste some links, and get traction. Someone needs to understand the community well enough to participate like a member rather than a marketer wearing a fake mustache.
The platform rewards usefulness.
At least when the system is working the way it is supposed to.
Measuring AI Visibility Is Still Messy
The measurement side gets complicated quickly. A company may be able to see AI bot visits in its server logs. It may see strange long-tail searches in Google Search Console that look like fan-out queries. Bing Webmaster Tools can expose more specific information for Bing and Copilot-related activity.
But none of that reconstructs the whole original user prompt.
Ruble made an important distinction between observing fan-out queries and knowing what a person originally asked. A company might see several unusual searches, but it may not know that those searches were all generated from one prompt. It also may not have enough timing detail to reliably group them together.
So claims about exact prompt volume deserve skepticism.
A lot of skepticism.
Ruble was especially critical of companies claiming to know search volume for AI prompts. He described two questionable ways that kind of data can be manufactured. One is to take traditional search-volume data and turn related queries into synthetic prompts. The other is to obtain real prompt data from network traffic gathered through free VPNs, browser extensions, coupon tools, or similar software.
The second approach creates an obvious privacy problem.
It also creates a sampling problem.
Even if the prompt data is real, the audience may not represent the company’s actual prospects. If the data disproportionately comes from people using free VPN products, is that really the population a mid-market ERP company should use to make marketing decisions?
Maybe not.
That is the kind of detail that gets lost when a dashboard produces a clean-looking number.
The Public Search Problem and the Private ERP Problem Are the Same Shape
This is where the episode made its most useful turn. Tim originally invited Ruble because he suspected that public AI discovery and private enterprise discovery were versions of the same problem.
They are.
On the public web, the AI interprets a question and retrieves information from search engines and web pages. Inside a company, Retrieval Augmented Generation (RAG) uses the same general pattern against private sources: understand the intent, determine what information is needed, retrieve it from available systems, and assemble an answer.
The scale and data sources are different.
The shape is very similar.
Tim had been hoping for something a little more magical on the private side. He wanted to think of ChatGPT as a kind of private Google Search for all of a company’s information.
But the conversation changed that model.
The LLM is not necessarily the search tool.
It is the intent tool.
And that is a big difference.
An LLM Does Not Automatically Make a Bad Search Tool Good
Once the LLM moves above the tool layer, it inherits the limitations of the tools below it. If the company’s document system has poor search, connecting an LLM to it does not automatically create perfect retrieval. If the ERP exposes data through an API, the LLM can call that API, but it is still constrained by what the API can expose and how the source system works.
For structured data, Tim summarized the idea in a way that felt almost disappointing. Maybe the AI is basically writing the SQL query that someone could have written before.
That is useful.
It is much easier for a normal user to ask a business question in English than to write SQL.
But it is not a fundamentally new database retrieval mechanism.
Ruble agreed with the framing that the LLM is a new intent mechanism. It can understand what the person wants, decide which underlying source is appropriate, create the query or tool call, retrieve the result, and turn it into a useful response.
That orchestration layer is powerful.
But it does not erase the architecture underneath.
MCP Gives the LLM a Standard Way to Use the Tools
Model Context Protocol (MCP) fits naturally into this picture. Ruble described MCP as a way to define how an LLM can use data sources and tools.
A company might expose Google Drive, SharePoint, a BI platform, an ERP system, or another application through an MCP connection.
The LLM can then decide which tool is most likely to contain the answer. For structured ERP or BI data, it might call an API and construct the necessary request. For documents, it might use the available document retrieval capability. The user does not necessarily need one giant unified search engine if the orchestration layer can choose among several tools.
That sounds like good news.
There is a catch.
The MCP connection still does not improve the underlying tool by magic. If a document repository has weak search, the LLM is still calling weak search. If an API has limitations, the LLM still has those limitations.
The AI is above the tools.
It is not replacing them.
That was probably Tim’s biggest architectural realization from the episode.
RAG Is Attractive Because Business Data Keeps Changing
The conversation also separated RAG from fine-tuning. A company can train or fine-tune a model using internal information, but that process represents a point in time. The business keeps moving after the tuning exercise is finished.
New orders arrive.
Procedures change.
Contracts get added.
Code changes.
Policies get updated.
That is one reason RAG is so attractive for business information. Rather than baking a static snapshot into the model and repeatedly redoing the tuning, the model can retrieve current information from the live source when the question is asked.
For ERP data, that distinction is huge.
Yesterday’s order data is not good enough.
The system needs the order that was entered five minutes ago.
Fine-tuning can still influence how a model behaves or which information it favors. But for constantly changing operational data, connecting the model to live sources is much easier to imagine maintaining.
Tim was looking for the “good news” on private data.
This is at least part of it.
The Prototype Is Easy; Production Security Is Not
Ruble added an important warning that deserves more attention than the cool demo usually gets. Giving an LLM access to company tools can create a side door into those tools if security is not designed correctly.
An MCP server with broad privileges is still broad privileges.
That is one of the differences between a prototype and a production system. It is easy to show an AI agent calling an ERP API and returning the right answer. It is harder to make sure that the user asking the question is only allowed to retrieve the data that user should actually see.
The question is not merely, “Can the AI get the answer?”
It is also, “Should this person be allowed to get this answer?”
For ERP systems, that is not a side issue. Financial data, payroll, customer information, vendor terms, pricing, and operational details already have security rules for a reason. The AI layer has to preserve those boundaries rather than route around them.
Very cool technology.
Still needs boring security.
Probably a lot of boring security.
Semantic Retrieval Moves Beyond Exact Keywords
Another piece that ties the public and private sides together is semantic retrieval. Ruble described the idea of representing words mathematically so that systems can recognize meaning, relationships, and similarity rather than relying only on exact text matches.
That is why an AI system can often match different wording to the same intent.
Traditional SEO trained everyone to think in keywords. If a particular phrase mattered, the instinct was to put that phrase on the page. With semantic retrieval, the exact phrase can matter less because the model can recognize related language and synonyms.
The objective shifts.
Do not write for the exact string.
Write for the intent.
Tim connected that back to the “good intent wins” idea. If AI is better at recognizing meaning across different wording, companies may have less reason to contort useful writing around awkward keyword formulas.
There are still technical ways to structure content well.
But good information has a better chance of being found even when the user asks for it in a different way.
That is encouraging.
The Original Content Matters More, Not Less
This also explains why Ruble is cautious about AI-generated “slop.” SpyGlasses can evaluate content against its model of the retrieval pipeline, identify problems, suggest changes, and even propose rewrites. But Ruble does not want the system to become another machine that generates thousands of generic articles from a keyword.
Why?
Because averaging the top ten existing articles mostly creates the average of what already exists.
Tim summarized the problem nicely: the goal should be to help someone stand out, not become the average of everyone else. Ruble said his company will provide outlines and optimization guidance, but the raw material still needs the company’s voice and knowledge.
That seems especially relevant in ERP.
A consultant, customer, developer, or software vendor usually has information that does not exist anywhere else. Implementation experience, unusual configuration decisions, benchmarks, real product behavior, troubleshooting history, opinions earned from actual projects, and original data are difficult to replace with generic generated copy.
Those are exactly the things worth putting on the web.
They are also exactly the things worth making retrievable inside the company.
AI Visibility Is Becoming an Engineering Problem as Much as a Marketing Problem
One thing Ruble finds interesting about AI visibility is that parts of the retrieval process are described in published research. He contrasted that with Google’s ranking algorithm, which is famously treated like a secret formula.
The AI systems are probabilistic.
But some of the research behind how those probabilities are calculated is public.
That is a funny reversal.
SpyGlasses tries to turn that research into something a marketing agency can act on. The software can run prompts, show which brands appear, identify cited pages and third-party sources, compare where competitors are cited, and score a page against a modeled citation pipeline.
Ruble was careful about the limitation.
They do not work for ChatGPT.
Their model is not perfect.
That caveat is important because this space changes constantly. The objective is not to claim perfect knowledge of a proprietary system. It is to use observable behavior and published techniques to create a better probability of being retrieved and cited.
That sounds less like classic copywriting and more like applied retrieval engineering.
Which is probably where this whole area is heading.
The Real Shift Is From Keywords to Intent
By the end of the conversation, the public marketing problem and the private ERP problem had converged on the same concept: intent. On the public side, the AI takes a long natural-language prompt and translates it into the searches and sources needed to answer it. On the private side, the AI takes a business question and translates it into the ERP query, document search, API call, or BI request needed to answer it.
That is what the LLM is exceptionally good at.
Understanding what the person means.
But intent still needs infrastructure underneath it. Search indexes matter. APIs matter. Web pages matter. Third-party sources matter. Structured data matters. Semantic models matter. Security matters. Server logs matter.
AI does not make those things irrelevant.
In some ways, it makes them more important because far more people can now use them without knowing the technical syntax.
So What Actually Changed?
Tim started the episode thinking AI discovery might require learning an entirely new world. He ended it with almost the opposite conclusion.
A lot of the old world is still there.
Search is still there.
Authority is still there.
Useful third-party references are still there.
Good data is still there.
APIs are still there.
Security is definitely still there.
The new layer is that a probabilistic language model can sit above all of those systems and translate human intent into the mechanics required to retrieve the right information. That is a big change, especially for people who never learned SQL, API syntax, search operators, or where every document lives.
But it is a layer.
Not magic.
And maybe that is the most useful way to think about AI in ERP right now. Instead of asking whether the LLM replaces the ERP, the BI tool, the search engine, the website, or the knowledge base, ask how well it can understand the user’s intent and route that intent into the right underlying system.
Then ask the uncomfortable second question:
How good is that underlying system at finding the answer?
Because the AI can only surface the right information at the right time if the information can actually be retrieved in the first place.
