Authors: Sam Koppelman, Michelle Cera, Matthew Termine
Editor: Vikas Kumar
Based on Hunterbrook Media’s reporting, at the time of publication Hunterbrook Capital is long $INOD. Positions may change at any time. This article is not investment advice or any recommendation. See full disclosures on our website.
Meta’s $200 billion bet on AI has already transformed an obscure New Jersey-based data company called Innodata. Hiring records and former employees reveal how closely the two are connected. Now, amid fears Meta might move away from Innodata after the tech giant’s investment in Scale AI, a new project focused on the “personalization of long-horizon agents” points toward Innodata’s involvement in Muse, which has been No. 1 in Apple’s app store for days.
Meta is Innodata’s largest customer, according to interviews with former employees, reviews of job postings, even an old short-report. In 2024, Innodata had roughly 400 full-time equivalent employees working on the Meta account, according to a former director. In the years since, that business has expanded.
In August, amid reported layoffs and fears that its Meta business was beginning to contract, Innodata disclosed a new program focused on “personalization of long-horizon agents” with its largest customer. At the time, it was unclear what this meant. Our reporting indicates that the project could be Muse, though neither company confirmed or denied the deal in response to repeated requests for comment. The size and materiality of this contract were not disclosed.
Hunterbrook Media routinely deploys open-source intelligence to identify tech deals and other secretive partnerships, before the companies involved post a press release. Past articles include Teradyne and Amazon, Datadog and Anthropic, Modine and Google, and Himax and TSMC. Email us at ideas@hntrbrk.com if you identify similar relationships based on publicly available information.
To oversimplify an extremely complex field of research, the two rate-limiting factors for the quality of frontier AI models are: 1) Compute; and 2) Data.
In a world where compute is constrained — where HBM has become the new Birkin; where Jensen has become a first-name-only celebrity; and where data-center protests are proliferating amid eschatological warnings of doom from the technology’s creators — the second part of that equation, the data, has become even more important.
The problem: AI companies have already ingested the low-hanging data. They’ve scraped vast stretches of the internet. They’ve bought and scanned (and allegedly stolen) books, to say nothing of news articles. Now they’re chasing ever more obscure data drops like teenagers on the SNKRS app.
Bro, I got the Spirit Airlines bankruptcy data, you can almost hear the Google procurement guy bragging, after scoring a treasure trove of emails, Teams messages, and spreadsheets last month. Limited edition. Distressed.
The backup bidder for that Spirit bankruptcy data, at $7.5 million, was Mercor, an AI-training startup.
But not everything an AI needs to learn is out there in some bankruptcy estate. Teaching a model to solve a hard problem — or carry out a novel or complicated task — can require fresh material: input from subject-matter experts, worked examples, and simulated situations in which it can practice getting things right.
That demand has helped make companies selling human expertise some of Silicon Valley’s hottest assets. Mercor went from a $250 million valuation in its 2024 Series A to $10 billion in its 2025 Series C — a 40-fold jump — with a rumored $20 billion round around the corner. Other companies in the space, like Turing and Invisible, have become unicorns, too.
Meta made perhaps the biggest bet with its $14.3 billion investment for a 49% stake in Scale AI, a major supplier of training data. The deal also brought Scale’s founder, Alexandr Wang, into Meta to help lead its AI effort.
After months of choppy headlines, Wang’s team finally has a consumer hit. A big one: Muse. Ten days after its launch this month, Muse reached No. 1 on Apple’s U.S. free iPhone app chart, with a 4.9 rating. Yesterday alone, Meta stock gained more than $100 billion (with a b!) of market cap.
Muse is marketed as a personal assistant for everyone. According to Meta, it can remember your preferences, work across your apps, use a browser to carry out tasks and keep going after you close the app. People are eating it up.
“The muse reception has honestly been beyond our biggest dreams,” Wang said on his very active X account, where he posted dozens of times in recent days highlighting the response to Muse.
“someone today said i had the best social team for the muse launch,” Wang boasted. “bruh this is all me out here rawdogging these tweets. i don’t let the corposlops near my tweet cannon.”
He later added a caveat:
Scale, however, is not the only data supplier with a place inside Meta. Another company has spent years helping train Meta’s AI. It isn’t a VC-backed startup. It trades on the Nasdaq. And its latest publicly disclosed assignment sounds an awful lot like the work required to make Muse useful.
But because it’s based in Ridgefield Park, New Jersey not Silicon Valley — and spent decades in obscurity — you probably haven’t heard of it.

It’s called Innodata ($INOD).
Founded in 1988, Innodata spent decades doing the unglamorous work of turning documents into digitized, structured information. Today, alongside that older business, it supplies training material, human judgment, and testing to help AI models improve.
That prospect helped turn a roughly $4 stock in early 2023 into a $121.50 stock at its June closing high. At the center of the transformation was a customer Innodata did not name in its financial statements. By 2025, that customer accounted for approximately $146 million — 58% of the company’s revenue. All available evidence indicates the mystery customer was Meta.
Innodata’s stock has started to slip in recent months, however, in part on fear that Meta might bring more of this kind of work in-house — after its investment in Scale. Why keep paying Innodata if you just bought almost half of a larger competitor? The second-quarter results added fuel to the fire: Innodata’s largest customer’s share of revenue fell by roughly a third, from an estimated $50.5 million to $34.1 million, using the company’s rounded customer-concentration disclosures.The insider selling from Innodata’s executives didn’t help, either. And Glassdoor is full of reviews reporting layoffs throughout 2026. The company’s shares are now back near $57.
Yet interviews with former employees, public hiring records, and the company’s own disclosures point to deep ties with Meta and an ongoing relationship. Innodata personnel have worked closely with Meta’s AI organization across multiple projects for many years, according to former employees. And Innodata’s new assignment from its largest customer — personalizing agents that perform long, complicated tasks — resembles some of the work behind Muse, which might only just be beginning.

In an interview with Hunterbrook Media, one former employee whose work covered new sales and service lines — and who was at Innodata after Meta’s investment in Scale — pushed back on the fear that Scale would displace Innodata. Asked whether Meta had started cutting its outside help after the Scale investment, they said: “If anything, I would say it’s more.”
Switching suppliers could introduce problems that would outweigh the savings, they argued. “What are they going to get? A 5% discount and headaches that they don’t know to anticipate.”
This is also the story Innodata has been telling investors. “With our largest customer, we continue to grow as we diversify into more organizations and more AI workflows and partner with them on their flagship next-generation AI program,” Innodata’s CEO Jack Abuhoff said on the company’s first-quarter 2026 earnings call. Innodata Chief Revenue Officer Rahul Singhal will take over as CEO effective September 30, 2026.
And Meta is no longer driving all of Innodata’s growth. Revenue from a second, unnamed “Big Tech customer” more than doubled in Q2, to roughly $31 million, almost matching the largest account. (It is unclear whether spending will continue at that level.)
Innodata also disclosed in August that it had won another new customer: “one of the fastest-scaling frontier labs.”
The education of a machine
What makes an AI model better: design or data?
In a recently published experiment, Dwarkesh Patel, perhaps AI’s preeminent podcast host, and researcher Jerry Han mixed models and training datasets from different years, pairing old with new to isolate the source of progress. It turned out, according to this study, that an old model trained with new data is more intelligent than a new model trained with old data.

The study involved small models, not Muse, and was limited in its scope, focused on pre-training. But its broader point was powerful and supports a longstanding hypothesis: Advances in training material can be a major source of advances in large language models.
This has been a tailwind for companies like Innodata, which focus on selling datasets to hyperscalers.
Innodata, in particular, has been focused on selling datasets designed for agentic use-cases. “Through our research efforts, we have established an early position in agentic reinforcement learning, one of the most important frontiers in AI development,” the company said in its August 6 press release. It also highlighted “a second program covering reinforcement-learning environments for computer-use agentic tasks.”
What, exactly, does this mean?
Innodata is now building training material for AI that acts instead of just answering questions. The company says its simulated environments teach agents to “reason, recover from errors, and complete complex, long-horizon tasks.” The environments come populated with fictional people, emails, documents and databases — and are deliberately complicated by missing information and conflicting instructions. Think of it as a practice version of the messy digital life an assistant will eventually be asked to manage.
Consider an instruction as simple as “plan my trip to Paris.”
An assistant like Muse might need to check your calendar, compare flights, find a hotel, arrange transportation, and ask for approval before spending your money. Along the way, it might notice that an old email lists dates you have since changed, remember your preference for nonstop flights, and distinguish the restaurant your friend recommended from the one they warned you to avoid. That is the kind of challenge behind Meta’s promise of an assistant that works across applications and remembers what matters to its user.
Some parts have checkable answers: Did the agent select the right dates? Stay within budget? Obtain permission? Innodata says it builds automated checks for those kinds of outcomes. But other decisions depend on judgment. Is saving $100 worth a six-hour layover? Does “somewhere lively” mean a neighborhood full of restaurants — or a room above a nightclub? A reservation can be perfectly valid and still be a terrible choice. The challenge is not merely teaching the assistant to book a trip. It is teaching it to understand what a good trip means to the person asking.
That is where human judgment becomes training data. Innodata says its reviewers score AI responses and explain their mistakes, following guidelines intended to “objectify subjectivity.”
In a travel exercise, a reviewer might favor an answer that respects the traveler’s stated priorities — or asks a necessary question — over one that confidently picks the cheapest option. Those comparisons can help developers teach models which responses to favor, rather than simply which facts to recite.
For a company like Meta, the potential value is an assistant that learns when to act, when to ask, and when mission accomplished is not quite the same as a job well done.
Whether this agent-focused business becomes a large, durable profit stream for Innodata remains to be seen.
For Innodata, the opportunity is not a royalty on every Muse download. It is the possibility that a product Meta keeps improving will require repeated rounds of training material, testing and human feedback. Innodata’s newly disclosed assignment appears closely aligned with that work. But the company has not disclosed the program’s size, duration or profitability, and new work could partly replace older assignments rather than add to them.
Innodata certainly still has questions to answer. One former employee who spoke to Hunterbrook likened Innodata to the taxi industry — up against more modern Silicon Valley counterparts they compared to Uber.
Short sellers, including Wolfpack Research and J Capital, have issued skeptical reports regarding Innodata’s AI capabilities and the size of its customer agreements. The company also faced investigations by both the Department of Justice and the SEC, but disclosed in June 2025 that both inquiries had been closed without enforcement action.
All of which is to say: Innodata is not without hair on it, especially during a year when company executives have sold shares hand over fist.
And the path ahead may not be a straight line. Management has left open the possibility of a quarter-over-quarter revenue deceleration. “Within the constraints of our business model, it’s certainly possible,” Innodata’s CEO said on its most recent earnings call.
But to the extent Innodata’s destiny is tied to the success of Meta’s AI platform, this has been a big week for the company — even if its involvement in Muse has yet to be confirmed … or denied.
Neither Innodata nor Meta answered Hunterbrook’s questions.
Muse, itself, was no more helpful when asked by a Hunterbrook reporter whether Innodata played any role in training it.
“If Meta used outside vendors on Muse, that would be a question for Meta’s privacy sources: the Privacy Center and the Muse Privacy Policy,” Muse wrote.
Muse added: “What the docs do say is that your conversations can be reviewed by Meta and used to improve AI unless you opt out in Settings → Data Controls.”
Whether those reviews are being led by Meta on its own — or involve outside vendors like Innodata too — remains to be seen.
Authors
Sam Koppelman is a New York Times best-selling author who has written books with former United States Attorney General Eric Holder and former United States Acting Solicitor General Neal Katyal. Sam has published in the New York Times, Washington Post, Boston Globe, Time Magazine, and other outlets. He has a BA in Government from Harvard, where he was named a John Harvard Scholar and wrote op-eds like “Shut Down Harvard Football,” which he tells us were great for his social life. Sam is based in New York City.
Michelle Cera trained as a sociologist specializing in digital ethnography and pedagogy. She completed her PhD in Sociology at New York University, building on her Bachelor of Arts degree with Highest Honors from the University of California, Berkeley. She has also served as a Workshop Coordinator at NYU’s Anthropology and Sociology Departments, fostering interdisciplinary collaboration and innovative research methodologies.
Matthew Termine is a former corporate lawyer with significant experience advising companies operating within regulated industries. Matt led Hunterbrook’s investigation and reporting on United Wholesale Mortgage. In 2017, Matt was credited by the Wall Street Journal, among others, for identifying suspicious mortgage loan transactions that led to several successful criminal prosecutions, including that of a prominent political operative and the chief executive officer of a federally chartered bank. He is a graduate of Trinity College and Fordham University School of Law.
Editor
Vikas Kumar joined Hunterbrook from The Capitol Forum, where he led the corporate investigations team for a decade as a senior editor. He was previously an attorney at Gordon Feinblatt, a trial attorney for the Department of Justice, and a law clerk for a federal judge. He has a J.D. from University of Virginia School of Law and a bachelor’s from Emory University. Vikas is based in Maryland.
LEGAL DISCLAIMER
© 2026 Hunterbrook Media LLC. When using this website, you acknowledge and accept that such usage is solely at your own discretion and risk.
Hunterbrook Media LLC (”Hunterbrook Media”) is an investigative news organization. Hunterbrook Media is affiliated with Hunterbrook Capital LP (”Hunterbrook Capital”), an exempt reporting adviser with the U.S. Securities and Exchange Commission that serves as investment adviser to one or more investment funds. Hunterbrook Media and Hunterbrook Capital are legally separate entities under common control. Hunterbrook Capital’s investment activities support Hunterbrook Media’s journalistic operations. Hunterbrook Capital’s investment performance can be affected by price movements in securities, derivatives, or other financial instruments related to companies covered in Hunterbrook Media’s reporting.
The specific position, if any, held by Hunterbrook Capital at the time of publication is disclosed at the top of this article. Any position or exposure may consist of direct holdings, short sales, options, swaps, other derivatives, or other forms of economic exposure to the securities or issuers discussed herein. Any position held may include equity securities, options, swaps, or other derivative instruments. Consistent with applicable policies and procedures, Hunterbrook Capital may establish, modify, or close positions in covered securities before or after this article is published. Following publication, Hunterbrook Capital will continue transacting in covered securities for an indefinite period and may hold long, short, or no positions at any time thereafter, regardless of any position described at publication. Hunterbrook Media has no obligation to update this article to reflect subsequent changes in Hunterbrook Capital’s positions. Any discussion in this article regarding value, valuation, downside, upside, or potential future stock price reflects Hunterbrook Media’s opinion as of the publication date. Such statements should not be interpreted as price targets and should not be understood to mean that Hunterbrook Capital intends to maintain any position until any particular price, valuation, or outcome is achieved. Hunterbrook Capital may modify, reduce, increase, hedge, or close positions at any time for risk management, portfolio management, liquidity, regulatory, investor, or other business reasons.
Nothing herein constitutes investment advice, a recommendation to buy, hold, or sell any security, or a solicitation to purchase or sell any securities. Hunterbrook Media is not a registered investment adviser in the United States or any other jurisdiction. All information and opinions presented are subject to change without notice.
Hunterbrook Media strives to ensure the accuracy and reliability of the information provided, drawing on sources believed to be trustworthy. Nevertheless, this information is provided “as is” without any guarantee of accuracy, timeliness, completeness, or fitness for any particular purpose. Hunterbrook Media does not guarantee the results obtained from the use of this information, and expressly disclaims any warranty, express or implied, as to accuracy or reliability. You should conduct your own research and seek advice from qualified financial, legal, and tax professionals before making any investment decisions based on information obtained from Hunterbrook Media.
The content provided by Hunterbrook Media does not constitute an offer to sell, nor a solicitation of an offer to purchase, any securities. No securities shall be offered or sold in any jurisdiction where such activities would be contrary to applicable securities laws. Hunterbrook Media authorizes redistribution of these materials, in whole or in part, for non-commercial, informational purposes only, provided that such redistribution includes this notice without alteration. Commercial use or alteration of these materials requires the express written approval of Hunterbrook Media LLC. By accessing this content, you agree to Hunterbrook Media’s Terms of Use.




