TL;DR — Key Takeaways

  • Enterprise data shouldn’t become a hidden second payment. Companies already pay for AI and software services—they shouldn’t unknowingly pay again with the proprietary data they contribute.
  • The real issue isn’t AI—it’s data ownership. Whether it’s training models, enriching databases, or building new products, vendors are increasingly looking to monetize customer data.
  • HubSpot became a cautionary tale. The company quickly reversed proposed data-sharing changes after customer backlash, underscoring the importance of transparency and explicit consent.

The recent discussion involving Palantir CEO Alex Karp and Microsoft CEO Satya Nadella raised an important question about enterprise data in the age of AI: Who should benefit when a company’s data makes somebody else’s model or product more valuable?

AI companies need data. The more relevant, proprietary and current that data is, the more useful the resulting models and services can become. Enterprise data is particularly valuable because it contains something the public internet cannot provide: the accumulated knowledge, decisions, relationships and experience of an operating business.

But enterprises are not simply providing data to AI companies. In most cases, they are also paying customers. They pay for the models, the cloud infrastructure, the tokens, the applications and the enterprise platforms through which AI is delivered. If the data they contribute then helps the provider build a better model or a more valuable product that it can sell to everyone else, the enterprise may be paying twice—once with money and again with data.

That is the concern underlying the Karp and Nadella discussion. But it is not merely an AI issue. AI has made enterprise data more valuable and intensified the competition to control it. It did not invent the business model.

HubSpot recently provided a particularly instructive example.

As Mi3 reported, HubSpot changed its terms as part of an initiative intended to support a shared data-enrichment capability. The company described its broader vision as “Trusted Prospecting,” an attempt to provide more accurate business-contact information, improve deliverability and make outbound prospecting more relevant.

On its face, that is not an unreasonable product objective. Anyone who has worked with CRM data knows how quickly business-contact information becomes inaccurate. People change jobs, email addresses stop working and once-valid records gradually become less useful. A continuously updated pool of professional contact information could create real value for HubSpot customers.

The problem was how HubSpot proposed creating that value and who would control the information used to create it.

HubSpot customers already pay the company to help manage their customer relationships. Those customers invest their own money, labor and institutional knowledge in building the records stored inside the platform. HubSpot then proposed using certain business contact information contributed through those customer systems to improve a shared enrichment product.

This was not primarily an AI-training controversy. HubSpot was not proposing to feed customers’ entire CRM histories into a large language model. It said the initiative concerned “business card-level” professional information rather than customer notes, deals, call recordings, custom fields or other sensitive CRM records.

That limitation is important, but it does not eliminate the underlying issue. HubSpot was still seeking to derive additional commercial value from information contributed by customers already paying for its software.

The connection to the AI debate is not about what technology was used. It is the economic relationship between the vendor, the customer and the customer’s data. Whether a company wants to use enterprise data to train a model, enrich a prospecting database, produce industry benchmarks or build its next commercial product, it is attempting to extract value from an asset created or supplied by a paying customer.

If I pay for a service, my data should not silently become a second form of payment.

HubSpot customers pushed back, particularly over the apparent use of default participation rather than a clear opt-in process. Four days after announcing the term changes, HubSpot withdrew them.

“We made a mistake,” HubSpot Chief Product and Technology Officer Duncan Lennox wrote in the company’s remarkably direct announcement reversing the decision.

HubSpot deserves credit for listening and responding quickly. Too many technology companies react to legitimate customer objections by publishing a longer explanation of why the customers failed to understand the original announcement. HubSpot did not do that. It apologized, withdrew the changes and said future enrichment capabilities that use customer data would be “fully and transparently opt-in.”

That was the right response. It also suggests the company understands what it missed the first time: The problem was not simply whether shared enrichment could produce a better product. The problem was who gets to decide whether the customer’s data becomes part of that product.

This is often treated as a privacy question. Privacy is certainly part of it, especially when systems contain personal, financial, medical or otherwise sensitive information. But privacy alone does not capture what is at stake.

This is also a question of economic ownership.

Enterprises spend enormous amounts of time and money creating the data held inside their technology platforms. Employees document customer interactions, maintain account records, correct inaccurate information, record product decisions and capture years of organizational knowledge. The software vendor supplies the platform that makes this work possible, and it is compensated through subscriptions, usage fees and service contracts.

The presence of customer data inside the vendor’s platform should not, by itself, give the vendor the right to repurpose that data for an unrelated commercial benefit. Hosting an asset is not the same as owning it.

There can be legitimate reasons for customers to contribute data to a shared system. Aggregated information can improve threat intelligence, fraud detection, benchmarking, product reliability and contact accuracy. In cybersecurity, every participating organization can benefit when signals from one customer help identify a threat targeting others. Similar network effects exist across many technology markets.

But that should be a deliberate exchange.

The customer should know what information is being used and for what purpose. Participation should be explicitly opt-in rather than enabled through a terms-of-service update that few customers will ever read. The customer should retain meaningful control over participation, including the ability to withdraw. Most importantly, the vendor should be able to explain what identifiable value the contributing customer receives in return.

That does not mean every customer contributing a data point needs to receive a royalty check. It means the exchange must be transparent enough for the customer to decide whether the benefit justifies the contribution.

Enterprise IT leaders therefore need to expand how they evaluate and negotiate technology agreements. Price, performance, availability, security and support remain essential. But buyers must also understand what the vendor is permitted to do with the data generated, stored or processed inside its platform.

Can that data be used to train a model? Can it be added to a shared enrichment service? Can the vendor aggregate it into a commercial benchmark? Can it be used to improve products sold to competitors? What happens to derived data after the customer leaves the platform? Does opting out disable a feature, or does it genuinely stop the vendor from using the data?

These questions belong in procurement reviews and contracts, not in an email sent after the agreement has already changed.

The pressure will only increase. As software functionality becomes easier to reproduce and AI models become more interchangeable, proprietary data becomes a more important competitive advantage. Every model provider, cloud platform, CRM company, productivity suite and enterprise software vendor will be tempted to find new value in the data flowing through its systems.

Some will create fair exchanges that benefit everyone involved. Others will decide that possession is close enough to permission and wait to see whether customers object.

HubSpot’s retreat was a victory for its customers, but it did not settle the larger issue. It merely showed what can happen when customers recognize that the bargain is being changed underneath them.

If a vendor wants to use my company’s data to improve its model, enrich its database or build another product, it can ask. It can explain the exchange, show me the value and let me decide.

But it should not charge me for the service and then quietly treat my data as a second payment. I already paid the bill.

Frequently Asked Questions

Why is HubSpot's policy change significant beyond CRM software?
It highlights a broader industry trend: software vendors increasingly see customer data as an asset that can improve their own products. The debate isn't about HubSpot specifically—it's about whether paying customers should also become suppliers of commercial value without explicit consent.
What's the difference between privacy and data ownership?
Privacy focuses on protecting sensitive information, while data ownership asks who benefits economically from enterprise data. Even if data isn't personally sensitive, companies should retain control over whether vendors can repurpose it for new products or services.
What should enterprise buyers ask vendors before signing AI or software contracts?
Organizations should ask whether their data can be used to train AI models, enrich shared datasets, create industry benchmarks, or improve products sold to other customers. They should also confirm whether participation is opt-in, how they can withdraw, and what happens to their data if they leave the platform.