Loading video...

Video Failed to Load

Go Home

niche business idea: laser graffiti removal > solo operator work > high margin / low overhead > sell recurring service contracts

636,014 views โ€ข 3 days ago โ€ขvia X (Twitter)

74 Comments

Nour Eddine Hamaidi's profile picture
Nour Eddine Hamaidi2 days ago

Doesn't have to be solo. You need the graffiti makers for ongoing work.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Better margins if youโ€™re solo- getting caught making the graffiti and getting paid to remove it would be pretty wild ๐Ÿ˜‚

Maitland's profile picture
Maitland2 days ago

@NOOROU it's not really *wrong* though... i mean some people like art, and some people like boring looking buildings... now they can have it both ways, AND you get to fire a laser blaster all day! AND some hoodlums get some extra drug money just for making their art! W's all around!

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

@NOOROU Wโ€™s all around. lol ๐Ÿ˜‚

Jonathan Challinger's profile picture
Jonathan Challinger2 days ago

Cons: - Doesn't remove green paint

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

haha, I think it does though. They were trying to show the contrast

The X-Philes's profile picture
The X-Philes2 days ago

Infinite money glitch: side hustle as a vandal at night

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Nighttime โ€œartistโ€ lol

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€3 days ago

Need the verdict from @CoFoundersNik and @mhp_guy Seems like this a good business if you can get the right contracts

charbwire's profile picture
charbwire2 days ago

I have 300-6000watt options for this solo operator.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

You run the laser engraving business?!

charbwire's profile picture
charbwire2 days ago

We are an SFX laser dealer. The USA branch is in western New York. We have a few companies that do this in the field. Itโ€™s becoming an interesting alternative or addition to surface prep companies.

OhioGuy's profile picture
OhioGuy2 days ago

@chasedownleads Guess who owns the spray paint vending machine around the corner?

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

@chasedownleads ๐Ÿ˜‚

OhioGuy's profile picture
OhioGuy2 days ago

@chasedownleads Infinite money glitch unlocked.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

@chasedownleads painter by night, cleaner by day

Yishan's profile picture
Yishan2 days ago

In our dystopian future, incentives eventually lead the business owners to fund local graffiti gangs

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

True. Thereโ€™s always perverse incentives in situations like this, but Iโ€™d hope a small business operator doing this in their hometown would have the best intentions. As long as a city contract was in place that didnโ€™t pay on a per-instance basis, there wouldnโ€™t be an incentive to create more graffiti.

J.D. Banker's profile picture
J.D. Banker2 days ago

@grok how much would the initial investment to get this business off the ground cost? What would the pricing model look like? Be as specific as possible.

Neil Hooey's profile picture
Neil Hooey2 days ago

A scene right out of Demolition Man.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Whoooooaaa youโ€™re right! I completely forgot about that scene That movie was ahead of its time

Neil Hooey's profile picture
Neil Hooey2 days ago

Yeah it predicted woke about as accurate as you could!

Norm Cruise Fan's profile picture
Norm Cruise Fan3 days ago

Let's see it not on fast forward and how it compares to other methods

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€3 days ago

Good point, the video is a tad fast. It might be possible to configure one on a rig/rack setup so that itโ€™s moving along a track fully autonomously The operator could take a pic and have AI asses the laser path for removal.

Norm Cruise Fan's profile picture
Norm Cruise Fan3 days ago

Sounds like a lot of work for what is currently a menial job performed at a low cost per m2. Usually you just paint over it, takes a few minutes

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€3 days ago

But for unpainted bricks, streets, and sidewalks it seems like a good idea It would be fun to be the townโ€™s laser guy :)

Norm Cruise Fan's profile picture
Norm Cruise Fan2 days ago

Yeah restoring original surfaces might be a winner there. Brick can be hella porous so prolly a bitch to scrub clean, but what do I know. Does the 'laser' blast away the top of the surface though? Might lose the finish. Couldn't bricks simply be sanded or water blasted?

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Not sure, but Iโ€™m assuming cleanup and transport might be easier with a laser rig. Plus they just look so fun :)

Norm Cruise Fan's profile picture
Norm Cruise Fan2 days ago

That's the fast forward playing tricks on you

Dy Mokomi's profile picture
Dy Mokomi2 days ago

This video is very much sped up. If you ever worked with these kind of lasers, you'd know that it's not suited for cleaning walls. Also it's a a very high powered laser that you can't just go a whip out in public

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Yep, it looks like a timelapse. As the technology progresses, Iโ€™m sure weโ€™ll see a much more powerful and portable version.

Dy Mokomi's profile picture
Dy Mokomi2 days ago

Unfortunately itโ€™s physics. Fiber lasers are already efficient. This can get more powerful (and a lot more dangerous) with bigger power source. Industrial fiber lasers are already at 6000W (obviously non-portable)

Degenerator's profile picture
Degenerator2 days ago

And then for marketing you just go around town doing a bunch of graffiti

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Gotta have a hobby ;) thereโ€™s actually enough in most big cities to keep you busy

Brian Sierakowski's profile picture
Brian Sierakowski3 days ago

looks like fun too :)

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€3 days ago

100% Iโ€™ve gone down a laser rabbit hole. Apparently, thereโ€™s a whole cottage industry for laser cleaning and removals Rust, barnacles, everything

niclas's profile picture
niclas3 days ago

Equipment costs 10.000-50.000 $

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

The fun is priceless :) haha A startup cost of $25k or so wouldnโ€™t be bad though. If you could secure a city service contract for $5k/month (or more) itโ€™d be worth it.

Turm Durmsen's profile picture
Turm Durmsen2 days ago

@erdverwachsener Takes way too long and doesn't work on every surface. Solvents or painting over is 100x easier/cheaper.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

@erdverwachsener not as cool or as fun as blasting away with a mega laser :) I could see this being done on sidewalks and parking lots too

Clanker Rights's profile picture
Clanker Rights2 days ago

then do graffiti on the side as an art form

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Hahah

แ—บye แ—บye Shorters ๐Ÿฆ‹'s profile picture
แ—บye แ—บye Shorters ๐Ÿฆ‹2 days ago

Step 4: create new graffiti at night to make more business

Adverse Selectee's profile picture
Adverse Selectee2 days ago

Niche business idea: laser execution of graffiti โ€œartistsโ€

"5th Generation Memes" Implied Carrot's profile picture
"5th Generation Memes" Implied Carrot2 days ago

It would be more effective to laser remove the ones doing the graffiti.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

True, but removing the โ€œartistsโ€ would limit the business potential ๐Ÿ˜‚

Jscott's profile picture
Jscott2 days ago

Downside your lings become a toxic waste dump for all the bullshit that burns off of the wall. Youd need a full hazmat suit with its own air to do that shit completely safely long term

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Upside: cool and profitable business idea

Jscott's profile picture
Jscott2 days ago

Id only do it if you know enough about chemical safety from burning materials and how to mitigate long term exposure, To the point youd bet your life on it Otherwise your profitable business is just you sitting around with lung cancer in 10 years, which is the opposite of cool

๐“›๐“ช๐“ญ๐”‚ ๐“๐“ต๐“ฒ Author,Editor,Voice's profile picture
๐“›๐“ช๐“ญ๐”‚ ๐“๐“ต๐“ฒ Author,Editor,Voice2 days ago

These are industrial Class 4 lasers, so you will also need proper safety gear, training, and usually local permits or insurance if you plan to operate them as a service. $16000-28000 is the startup cost if you don't already have a capable vehicle, most likely.

huebet's profile picture
huebet2 days ago

Uninsurable hazmat exposure, heavily regulated, prohibitive tool overhead Compare with solvent coat and pressure washer, no brainer

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Iโ€™ll bet you could get insurance and find a way to safely remove Tool overhead isnโ€™t much

Gonza Shwetz's profile picture
Gonza Shwetz2 days ago

Niche business idea: hitman who hunts down graffiti artists

Steve Coulier's profile picture
Steve Coulier2 days ago

This works especially well if you're also a graffiti artist.

Blazingsolos's profile picture
Blazingsolos2 days ago

The humble photon reflecting off a nail or a fence into the eyeball of a bystander

goldmantaysachs's profile picture
goldmantaysachs2 days ago

Can this technology be adapted to use on criminals?

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

What do you mean? Like tattoo removal or something? Probably too dangerous

goldmantaysachs's profile picture
goldmantaysachs2 days ago

I was thinking criminal removal

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Ahhh ok

silent clamor's profile picture
silent clamor2 days ago

@grok does this tech actually work for graffiti removal at scale? Is it super expensive? Why don't all cities use this all the time?

Amonte's profile picture
Amonte2 days ago

Europe needs this badly

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Very true, especially parts of Italy like Rome, Florence, Sorrento, etc.

Kine Dareel's profile picture
Kine Dareel2 days ago

High overhead. Places with graffiti is generally high crime. Some gangbanger sees you using some fancy expensive layer getting rid of their territory marking. Youโ€™re getting mugged, you lose your 10k equipment and your safety. Hope you had good insurance.

Vivien Reid-tardo ๐Ÿ‡บ๐Ÿ‡ฒ's profile picture
Vivien Reid-tardo ๐Ÿ‡บ๐Ÿ‡ฒ2 days ago

I need this to remove paint from a stone fireplace.

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Small home renovations would be super cool with the laser tool. Iโ€™ve found videos of them removing rust as well

chadutanunโ˜ฎ๏ธImagineโ˜ฎ๏ธ's profile picture
chadutanunโ˜ฎ๏ธImagineโ˜ฎ๏ธ2 days ago

Toxic fumes?

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

Yeah, probably not too healthy although there might be a way to safely extract the fumes away

Mananaba ๐Ÿฅž๐Ÿ—'s profile picture
Mananaba ๐Ÿฅž๐Ÿ—2 days ago

Just add a code to the paint and sell with name id

๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€'s profile picture
๐’Ÿโ„ฏ๐“‡โ„ฏ๐“€2 days ago

What do you mean by a code to the paint? Are you suggesting a way to identify paint purchases?

Mananaba ๐Ÿฅž๐Ÿ—'s profile picture
Mananaba ๐Ÿฅž๐Ÿ—2 days ago

Yes like different ratios of additive that will allow you to identify the created batch id. And selling them only with id. No anonymity... No graffiti

Robert Allison's profile picture
Robert Allison2 days ago

Nice! - I wouldn't mind running a business doing this!

Orcus's profile picture
Orcus2 days ago

Cons: Can only work between 5 AM - 10 AM to reduce the chance of getting shot by bangers.

ZEP's profile picture
ZEP2 days ago

I don't think anyone hiring for this cares about the method.

๐Ÿšš ๐Ÿ“ฆ โœˆ๏ธ๐ŸŽธ๐ŸŽง๐ŸŽ™๐Ÿ’พ๐Ÿ•ธ๏ธ's profile picture
๐Ÿšš ๐Ÿ“ฆ โœˆ๏ธ๐ŸŽธ๐ŸŽง๐ŸŽ™๐Ÿ’พ๐Ÿ•ธ๏ธ2 days ago

The cool thing is that you can also do laser hair removal between jobs.

Related Videos

๐Ÿšจ THE NISSAN ALTIMA BECAME AMERICAโ€™S โ€œLOW CREDITโ€ CAR โ€” AND WHAT NISSAN DID HAS PEOPLE STUNNED A breakdown going viral claims Nissan didnโ€™t just sell carsโ€ฆ it built a system around high-risk loans. Not performance. Not reliability. Debt. โ€ข Easier approvals than competitors โ€ข Interest rates reportedly pushed past 20% โ€ข Loans stretched to 7 YEARS to keep monthly payments looking โ€œnormalโ€ โ€ข The higher the interest rateโ€ฆ the MORE that loan could be sold for โ€ข Loans bundled and flipped to investors for $4,000+ profit per deal โ€ข The car becomes secondaryโ€ฆ the loan becomes the real product Buyers with strong credit shop for the lowest rate. Buyers with weak credit take what they can get approved for. That gap? Thatโ€™s where the money is. And over time, one car kept showing up in those deals with repossessions, missed payments, โ€œlast resortโ€ financing... The Altima. But it didnโ€™t come free. โ€ข Higher repossessions โ€ข Owners struggling to maintain the cars โ€ข A stigma that stuck โ€ข Sales and brand perception taking a hit The strangest part? The same company behind the Nissan GT-R went from performance iconโ€ฆ to being tied to one of the most controversial lending reputations in the market. Is this smart businessโ€ฆ or a system that made more money the worse people did? ๐Ÿ“น: rohrsteam

HustleBitch

39,615 views โ€ข 4 months ago

My last post about starting a stump removal business went viral, and my DMs filled up with the same question: how do I actually get started in this business? This is a great blue collar hedge against AI if you are worried about losing your white collar job in tech or otherwise. I've made $12,000 in the first 3 months and here's everything I have learned in that time: The machine is the big check. Everything else is small. I bought a Bandit SG-40 tracked stump grinder for $33k and financed it at about $1k/month. The catch is the interest, roughly $100/month at a high rate, so I'm just paying it off next month to make the interest go away. I also have a $1M general liability policy that costs about $100 per month. Plus the finance company made me get insurance on the machine for about $100 per month. You need a way to haul it. That's a trailer, not a new truck. I got a Big Tex trailer for $4,000. I already had a 4Runner, so I didn't need a truck to start. I'm buying an F-350 diesel with the 7.3L for $6k, but only because a close friend is selling it to me and that's an insane deal. Do not go buy a $60k truck to start a stump business. Haul it with something cheap if you have to buy a truck just to get things started. Pricing: $5โ€“$8 per inch, and you feel it out. I charge $5 to $8 per diameter inch and I read each job. A stump sitting high off the ground or one they want ground deep costs more. A flush little stump in soft dirt costs less. Measure the diameter, pick your per-inch number based on difficulty. And set a minimum. Mine is $200 for me to even show up. Below that it's not worth the drive and setup. Referrals from tree services are the whole game. Tree companies cut trees down. They usually don't grind the stump. That leftover stump is your job, and they're standing right next to it. Go build relationships with every tree service in your area. When they hand you their stumps, you have a pipeline you didn't have to advertise for. But try every channel, because they all work. I've landed a paying job from every single marketing channel I've tried so far. Not one dud yet: Flyers on community boards Business cards I hand out locally at the cafe and the dog park Posts in the local Facebook groups Nextdoor Yelp Google Business Profile Corrugated cardboard road signs None of it is fancy. It's just showing up where people already look when they've got a stump in the yard. The next one I'm testing: post office mailers. You make a flyer, pick a route, and the post office lets you see the demographics of that route, age, household income, how many people live there. Then it's about 27 cents to mail each one. You aim it at neighborhoods that own homes with yards. I'll report back on how it does. Support local, and local supports you. Use the local businesses. Buy from them, show up, be a regular. They turn around and send business your way. A small town runs on who knows you, and stump removal is a local business. Yes, it's a one-man job. Yesterday I ran two jobs, made $300 and $375. From leaving my house to pulling back in was about 4 hours, and I burned maybe $20 in fuel between my car and the machine. That's a good day when you've got two small ones. Three months in I've made $12,000, averaging about $4k/month. I expect that to climb as more people in the area know my name. The math is simple: low overhead, cash jobs, work that has to be done in the real world. Bonus: this is a business AI can't take from you. I spent 10 years in an office. With everything up in the air around AI and jobs, there's something I like about owning a business where the work physically has to happen. A stump doesn't grind itself. You show up, it's there, you leave, it's gone. And after a decade at a desk, I forgot how good it feels to see exactly what you accomplished at the end of the day. I've shared a vid of a job I did yesterday ๐Ÿ‘‡

smokey

34,218 views โ€ข 1 month ago

Warren Buffett literally gave a 9-minute masterclass on what makes a business worth owning, inside the interview where he explains why he broke his own rule on technology. Eight things he teaches: 1. A good business is not one that grows. It is one that earns high returns on capital for a long time. His words: "something that you can expect to earn high returns on capital over a long period of time." Growth without returns on capital is just a bigger version of the same problem. 2. Measure it against doing nothing. Buffett points out he can put huge amounts of money into government bonds and collect payments every year with no risk. So a good business has to earn a lot more than treasuries, and be expected to keep doing it. If your business does not clear the riskless rate by a wide margin, the capital has a better home. 3. The gap between similar-looking businesses is enormous. Most banks earn 13 or 14 percent on capital. Ask anyone to guess American Express and they say something similar. It earns 30 percent plus, and Buffett is clear it "does not incur more risk in doing so than the banks that earn 13 or 14 percent." Same industry, more than double the return, no extra risk taken. 4. Charlie Munger's test: the cash has to be real. Munger pounded the idea that a business was not good just because it was doing sexy things. It had to be earning real cash, be able to pay that cash out if it wanted, and better yet be able to put it back to work inside the business. A company that earns high returns but cannot redeploy the money is worth less than one that can. 5. Time is the multiplier, so duration is the thing to protect. Buffett says a long period of time "gets to be very important because it doubles later on to the very big numbers." One great year is noise. The rate is what compounds. 6. When the facts change, retire the rule. Buffett spent decades known for not buying technology, and said so himself. His explanation for buying now is that the business changed: Google and its competitors are "laying out hundreds of billions," they are big capital spenders, and that is real money. When they were asset-light he passed and the market loved them. Now that they spend heavily, shareholders like them less and he thinks they are more likely to win. He did not change his test. He noticed the business had moved into the category his test rewards. 7. Nobody is measuring the thing that matters. Buffett says he cannot recall a report on Wall Street that gets into the internal rates of return a business is actually earning, and calls the fixation on next quarter ridiculous. He rates Alphabet ahead of 90 or 95 percent of what gets merchandised through Wall Street, on the record rather than the story. If your own reporting tracks growth and headcount but not return on capital, you are measuring what is easy. 8. Every wonderful business gets attacked, so ask how long it stays wonderful. In 1958 he helped start Data Documents, after IBM was forced by an antitrust settlement to divest half the capacity of its best business. That advantage ran out after 10 or 15 years, and he knew some of the people who caused it to run out. His closing line is the whole lesson: "It's not a question of whether it was wonderful yesterday. The question is, how long is it going to be wonderful?" The move for an operator: run the test on your own business this quarter. What return are you earning on the capital in it, how does that compare to doing nothing, and what would have to be true for that return to survive the next ten years. Warren Buffett with Becky Quick, CNBC Squawk Box, July 2026.

Andrej Drats

31,661 views โ€ข 2 months ago

In 2026, Venture Capital will eat Private Equity It used to be that venture capital and private equity lived on two separate planets: VC = San Francisco PE = New York They targeted completely different universes of companies: --> PE - people heavy biz services, stable/low growth, predictable cashflows --> VC - tech-forward, high growth, high risk, massive TAM What was the playbook for B2B VC backed startups? --> Grow to unicorn scale by selling to other early adopter tech companies, then Fortune 500s XX> SMB and mid-market services - think field services, IT staffing, accounting, construction, recruiting - were always tough to sell into for startups Why? -->Thin margins, high labor costs, and small IT budgets >> But as AI eats labor, these businesses are in play << There are 3 ways where VC and PE are colliding: 1/ Private Equity funds will become channel partners for startups. PE funds are focused on financial engineering and cost optimization. Startups building AI products and services can sell across their portfolio to automate the backoffice and uplevel sales and marketing. PE funds have made AI their #1 strategic priority and have hired central leaders to oversee their portfolio adoption efforts 2/ PE portfolio pages are a startup idea menu Private equity will often buyout vertical software companies whose TAM didnโ€™t allow venture scaled returns. As software evolves from data storage and collaboration to agents taking action and completing work, AI should massively expand the TAM for these categories. Founders will set their sights on unseating these legacy incumbents backed by private equity. All they have to do is look at their portfolio pages for category ideas 3/ AI Rollups This is one of the most direct ways that VC is eating PE VC backed AI platform businesses are not just selling software but acquiring legacy business services companies to own the value chain end to end. As an example, our speedrun company AgentAstra is acquiring freight forwarding services businesses with mostly debt and integrating AI deeply into their operations These companies aim to increase margins by at least 2x and make them โ€œAI nativeโ€ tl;dr - While the west coast, Patagonia-wearing VCs and the east coast, PE suits used to live in different universes, in 2026 with AI, I believe, those worlds converge

Troy Kirwin

187,377 views โ€ข 9 months ago

This is part 2 of a 2 part post (see part 1 here Below is a structured analysis to demonstrate the validity of using buyers of Veritaseum #SmartMetal to buy into and sell compute from globally aggregated cell phone compute pools - directly compeiting with the big guys - Google, Amazon and Microsoft cloud businesses. We discuss estimates, business model propositions, and potential economic outcomes, but first, see my Executive Global Article on Zero Profit Models( and purchase Veritaseum SmartMetal here - Can you really disintermediate the most profitable revnues of t $6.7 trillion worth of technology cloud providers? Well, the fact that it is among, if not the, most profitable of their revenue drivers is a very material clue! Step 1: Estimating the Number of High-End Smartphones Globally As of early 2025, approximately 7.5 billion smartphones are actively used worldwide. Considering that: About 30% of global smartphones are high-end (comparable or superior to an iPhone X; for instance, Samsung Galaxy S22/S23 Ultra, iPhone 16 Pro Max with A18 chips, and Qualcomm Snapdragon 8 Gen 3 or newer). Thus, approximately 2.25 billion high-end smartphones exist today (30% of 7.5B). Step 2: Aggregate Compute Power Estimation (Idle Capacity) Average Computational Capacity per High-End Smartphone: A high-end phone has roughly: CPU: ~1 to 1.5 TFLOPS GPU: ~1.5 to 2 TFLOPS Average Idle Compute per Phone: 1 TFLOPS (CPU) + 1.5 TFLOPS (GPU) = ~2.5 TFLOPS idle. Total Potential Compute Power: 2.25B smartphones ร— 2.5 TFLOPS each โ‰ˆ 5,625,000,000 TFLOPS (5.625 ExaFLOPS) Comparison to Cloud Vendors: Amazon AWS, Microsoft Azure, Google Cloud combined currently deploy approximately ~1 to 2 ExaFLOPS of continuous computing power. Thus, aggregate idle compute power from high-end smartphones (5.625 ExaFLOPS) exceeds the largest cloud vendors combined by at least 2.8x. Step 3: Proposed Business Model ("Zero Margin Trustless Model") Following Middletonโ€™s economic principles, a decentralized marketplace based on his IP (SmartMetal Rounds and patented protocols) would allow individual users to rent their smartphonesโ€™ idle compute power. The economics would follow: Revenue Structure: Compute resources provided by phone owners (children, elderly, economically disadvantaged communities) rented to consumers (AI firms, universities, research institutions, enterprises). Offered at 10% above net cost ("as close to free as possible" per the attached articleโ€‹Executive Global articlโ€ฆ). Revenue Distribution: SmartMetal Owners (phone owners): Receive 20% of net revenue generated. Platform Cost & Overhead: Costs for electricity, network management, and maintenance (approximately 70% of net revenue). Intellectual Property Licensing (Middletonโ€™s IP): A modest licensing feeโ€”around 10% (aligned with Middletonโ€™s zero-margin, IP-licensing-centric model). Step 4: Revenue Estimation Example Assumptions: Average monthly idle compute contribution per phone: 4 hours/day, 30 days = 120 hours/month. Market price for decentralized high-performance computing: approximately $0.10 per TFLOP-hour. Revenue per Smartphone per Month: Compute provided: 2.5 TFLOPS ร— 120 hrs = 300 TFLOP-hours Revenue at $0.10 per TFLOP-hour: 300 ร— $0.10 = $30/month per smartphone Aggregate Monthly and Annual Revenue: Monthly revenue (2.25 billion phones): $30 ร— 2.25B โ‰ˆ $67.5 billion Annual revenue potential: $67.5B ร— 12 months = $810 billion annually Distribution of Annual Revenue: SmartMetal Round Owners (20%): $810B ร— 20% โ‰ˆ $162 billion/year Operational Cost (70%): $810B ร— 70% โ‰ˆ $567 billion/year Middleton IP Licensing (10%): $810B ร— 10% โ‰ˆ $81 billion/year Thus, the total economic benefit is substantial, particularly transformative for economically disadvantaged participants (children, elderly, developing regions). Step 5: Practical Impact & Social Value Impact on Children & Young Adults: Empowerment through earning potential (around $360 annually per child smartphone owner). Practical, intuitive introduction to economics, technology, and entrepreneurship through gamified interfaces and secure, decentralized platforms. Impact on Elderly and Economically Disadvantaged Communities: Significant supplemental income (potentially exceeding many pension plans or assistance programs). Bridging the technology gap, ensuring inclusive participation in global digital economies. Step 6: Strategic Value & Market Positioning Middleton's patented Zero Margin Trustless Model ("ZMTM")โ€‹Executive Global articlโ€ฆ creates a highly attractive, low-cost computational offering. Competing directly with incumbent cloud providers: The computational marketplace can massively disrupt cloud computing with lower fees and broader global reach. Leveraging Middletonโ€™s IP and SmartMetal Rounds, it creates defensible competitive barriers and immense value for early adopters. Step 7: Driving Middletonโ€™s Peer-to-Peer Economy As described in Middletonโ€™s visionโ€‹Executive Global articlโ€ฆ, this marketplace underpins a global peer-to-peer economy, transforming idle smartphone resources into meaningful economic output. The P2P economy will leverage: AI-driven autonomous economic agents. Secure blockchain-based IP rights enforcement. Economic democratization by redistributing traditional cloud revenues directly to everyday device owners. Summary & Strategic Conclusion Implementing a decentralized compute platform powered by high-end smartphones and Middletonโ€™s patented Zero Margin Trustless Model presents enormous economic potential, far exceeding current major cloud vendors combined. With annual revenues estimated up to $810 billion, and meaningful income distribution to disadvantaged demographics, this innovative model could dramatically reshape the global computational economy, achieve significant social impacts, and provide the backbone for Middletonโ€™s envisioned peer-to-peer decentralized economy.

Reggie Middleton, Disruptor-in-Chief

14,751 views โ€ข 1 year ago

Long $MCB.TO โ€” oilfield automation and services, $90mm CAD market cap. This is a 111 year old company thatโ€™s been providing oilfield equipment for over 50 years. The current CEO has been with the company since 2002. Profitable, growing, trading at 9x TTM earnings backing out net cash, and beginning to introduce a pretty revolutionary full product suite to automate casing running - the process of lowering sections of steel casing into an open wellbore to act as a liner, reinforcing the wellbore and acting as a channel through which oil and gas can be extracted. This is a pretty critical step in drilling any new well, and has applications in both onshore and offshore oil and gas. If you guys have seen those viral videos of oilfield workers running casing, you know how dangerous and physically demanding this job looks. Itโ€™s extremely tough stuff โ€” comes with a lot of injuries, itโ€™s very difficult to hire for, and because the work sucks so much and nobody wants to do it, the pay needs to be fantastic. Itโ€™s a painful and expensive process for operators and oilfield workers alike, see the attached content. $MCB.TO has spent 6 years developing a product suite that can effectively automate this process, more than halving the usual need for skilled workers, with labour costs typically making up 50% of casing running costs. $MCB.TO โ€˜s automated product suite increases safety, efficiency, and cuts costs significantly for oilfield operators to complete this process. Their automated product suite here includes a SaaS package, without which none of the other equipment works. Their software facilities communication between all the necessary pieces of equipment, monitors torque levels, analyzes casing rotation speed, and it includes a sensor suite that can monitor for various extreme conditions during a casing run โ€” all of this automation significantly decreases the workload for casing crews and eliminates a significant chunk of the involved labor. It can also ensure higher wellbore integrity. Theyโ€™re first to market with this stuff โ€” competitors have bits and pieces of separate equipment, but theyโ€™ve got the full package, and itโ€™s a much better offering. It would likely take a similar amount of time and investment for competitors to copy this suite, and some of their new high-tech equipment is specifically IP protected. The margins on their software system will be sweet, much higher than current mid-30s gross margin. So youโ€™ve got high incremental margin expansion baked into every product suite these guys sell, and theyโ€™re going to see significant topline growth given the scale of the problem theyโ€™re addressing and the willingness with which operators are going to integrate this tech. The existing business is already solid. No debt, $10mm CAD cash position, enough working capital to fund their new system. Theyโ€™ve stayed in business for over a century, Iโ€™m sure they know what theyโ€™re doing. Current revenues are a 50/50 split between Americas and ROW, shipping to 50 countries any given year. Theyโ€™ve already got a 40% recurring/repeat revenue split selling oilfield consumables, replacement products, and servicing equipment. Theyโ€™ve already begun rolling out their new product suite and receiving SaaS contracts. Iโ€™m expecting them to see significant topline growth from their new products with a lot of margin expansion to boot. They pay a 3% dividend and are buying back shares here. Dwindling rig counts are a risk. I canโ€™t forecast where the price of oil will go or how the drilling process will shake out, but in previous years they have increased sales despite dwindling rig counts simply because their technology is a better solution โ€” I donโ€™t see why they couldnโ€™t do that again. When oil prices eventually head back up, rig counts will follow suit, and many oil operators will likely go straight to $MCB.TO โ€˜s product suite. Still some more research I need to do but I think these guys will do well. Long 12% allocation @ $3.26 CAD.

Welfare Capital (Jack)

23,541 views โ€ข 1 year ago

The majority of neoclouds will eventually go out of business but here is the winning formula if you want to win. (Save this). The core problem for the industry is that the economics of running GPU infrastructure only work at massive scale, with cheap financing and investment grade customers backing long term contracts. A lot of the names crowding the middle column of that chart are Bitcoin miners who converted their rigs into GPU racks chasing the AI trend, rather than companies built from the ground up for this business, which is exactly the kind of opportunistic entrant that gets wiped out when capital tightens or utilization dips. Nebius sits in the Neocloud Giants tier alongside CoreWeave, Lambda and Crusoe, and today's Q2 2026 print showed exactly why it's pulling away from the pack rather than getting lumped in with the 78 emerging players facing consolidation risk. Revenue hit $582 million, up 454% year over year, with annualized recurring revenue reaching $3.0 billion by the end of June, up 58% quarter over quarter. The company won four separate customer agreements each worth over $1 billion in total contract value and total contract value won during the quarter jumped 4x versus the prior period, a growth rate most of the smaller neoclouds on that chart simply can't match without hyperscaler grade balance sheets. Here's the vertical stack that sets Nebius apart from most names on that chart. Unlike pure GPU rental shops that lease space in someone else's data center, Nebius designs its own data centers, builds its own server racks and motherboards, procures its own compute, and runs a proprietary AI specific cloud platform layer on top of all of it. That full stack control, from silicon to software, is precisely what most of the emerging neoclouds in the chart's middle column lack, since converting a Bitcoin mining facility gives you power and cooling, but not in house rack engineering or a purpose built cloud software layer. Nebius has also been shifting from leased to owned infrastructure, with more than 75% of its contracted power now sitting at facilities it directly controls, up sharply from a mostly leased model just a year ago. That ownership shift is the difference between capturing margin over the long run versus being at the mercy of a landlord's lease terms, which is a structural advantage over neoclouds still renting third party space. Now for the pricing power piece. Nebius disclosed today that it's now charging $40-50 million per megawatt on new capacity deals, already signing its first one at that price this week, up from roughly $12 million per megawatt on its 2026 base contracts. Management also said it could sell its entire 2027 capacity right now on these terms but is deliberately holding some back for near term customer needs, a level of pricing leverage that tells you demand is outstripping supply for anyone offering real scale and reliability, exactly the customer profile the smaller, undercapitalized neoclouds struggle to attract. Nebius isn't a single product company either, which matters given how many names on that chart have no other legs to stand on if GPU rental margins compress. The company owns Avride, an autonomous driving and delivery robotics business with partnerships with Uber and Hyundai, TripleTen, a tech re skilling edtech platform and holds equity stakes in ClickHouse, the database company it spun out and recently backed in a funding round, and in Toloka, an AI data-labeling platform that sold a majority stake to Bezos Expeditions and Shopify in 2025. Those side businesses give Nebius optionality and diversified cash flow that a converted mining rig operator simply doesn't have. Bullish on Nebius, make sure to follow Melvin for more AI infrastructure insights, and if you want to see exactly what I'm buying as an analyst at Milk Road Pro, you can check out the link below for more.

Melvin

37,428 views โ€ข 1 month ago

$SIVE Sivers Semiconductors: The Photonics Inflection In the semiconductor world, real alpha is found where physics hits a wall. Today, that wall isnโ€™t GPU compute power - itโ€™s interconnect bandwidth. As we transition to 1.6T networking, copper is dying, and light is taking over. Sivers Semiconductors ($SIVE) is no longer just a "Swedish tech hope." It has officially transitioned from an engineering research house to a high-volume product company. 1โƒฃ The 1.6T AI Bottleneck: Indium Phosphide (InP) AI clusters are only as fast as the links between them. Silicon Photonics (SiPh) is the solution, but silicon cannot emit light efficiently. It needs an external "engine." โžก๏ธThe Moat: Sivers is one of the few global players capable of mass-producing InP CW-WDM laser arrays. These are the "spark plugs" for the next generation of AI transceivers. โžก๏ธProof of Concept: Partnership with $POET is hitting a critical milestone. Prototype External Light Source (ELS) modules for 1.6T architectures are sampling in H1 2026. โžก๏ธThe Pivot to "Standard Products": CEO Vikram Vathulya recently confirmed a strategic shift. Sivers is moving away from low-margin custom engineering toward Standard Products. This will drastically shorten "time-to-revenue" and scale margins by serving multiple customers with the same high-spec chips. 2โƒฃ Hard Evidence: The 2026 Contract Ramp-up Investors have long criticized Sivers for a "paper pipeline." That changed this month (March 2026): โžก๏ธLiDAR Breakthrough: A strategic LiDAR customer (winning in both Automotive and Industrial) is ramping up in Q4 2026. Cumulative revenue potential: $53M to $138M. โžก๏ธSATCOM & IRISยฒ Momentum: The Wireless division grew 33% in 2025 (constant FX). Crucially, three terminal vendors for Europe's IRISยฒ satellite constellation have moved to the RFP stage and are currently building prototypes using Sivers technology. โžก๏ธUS Chips Act: Sivers is using Chips Act funding not just for cash, but to accelerate the integration of their tech into US Defense "Electronic Warfare" (EW) programs. 3โƒฃ Financial De-Risking & The "Uplisting" Catalyst The biggest drag on $SIVE has been its balance sheet. That drag is being cut: โžก๏ธDebt Refinancing (Feb 2026): Secured a $17M facility from Bootstrap Europe, consolidating all debt and providing a clear runway to the Q4 2026 ramp-up. โžก๏ธThe 2027 Line in the Sand: Management has set a firm target to reach full break-even/positive cash flow by the end of 2027. โžก๏ธThe US Nasdaq Spin-off: With 80% of Photonics revenue coming from the US, the plan to spin off Sivers Photonics into a US-listed entity remains the primary "valuation unlock" to capture US-style multiples (think Lumentum or Coherent). 4โƒฃ 2026 Guidance: The Roadmap to Pavement โžก๏ธOpportunity Pipeline: Stands at $453M (up 64% YoY). โžก๏ธProfitability Pivot: Q4 2025 delivered a positive Adjusted EBITDA of $1.14M. Expect this to stabilize as "Foundry Customers" (SME base business) provide a recurring revenue floor while waiting for the "Big Elephants" (AI & Auto) to join. โžก๏ธOFC Los Angeles (March 15-19, 2026): Currently underway. Industry leaders are vetting Sivers' laser arrays. Success here is the catalyst for large-scale datacenter deployment. ๐Ÿ‘‡Final Verdict Sivers is no longer a "story" stock; it is a "delivery" stock. As 1.6T networking becomes the standard for AI datacenters, the demand for Indium Phosphide laser sources is set to explode. Sivers is one of the very few companies sitting on the right IP at exactly the right time. Whatโ€™s your take on the Silicon Photonics race? Are you betting on the massive, vertically integrated giants like Broadcom, or do you see the "pick-and-shovel" specialists like $SIVE capturing the real alpha in the 1.6T transition? Drop a comment below with your thoughts or ask me anything. I'm here for you. #Investing #Semiconductors #AIInfrastructure #StockPicking #Sivers #Photonics

Finn Stockinger

351,610 views โ€ข 6 months ago

๐ŸŸ My largest single name equity position since $XOM in 2020โ€ฆ $GEO I have been accumulating a position for a few months. Some people I know are already familiar with the name, many are not. Michael Burry actually made headlines a year or two ago for initiating a position himself. The Geo Group is a private prison/detention facility operator that also provides โ€œATDโ€ (Alternative to Detention) services such as GPS tracking devices and electronic monitoring. In the US alone they operate 50 secure facilities and 65,000 beds. They have contracts with the US Marshals and US Immigration and Customs (ICE). GEO is well positioned to benefit from the current immigration crisis, which I personally believe is the worst domestic crisis the US has faced in many years. Itโ€™s the leader and has market dominance in the Alternative to Detention space, with existing government contracts already. Their ATD solutions for illegal migrants/asylum seekers are more humane, cheaper (~25x cheaper to monitor per day than detain), and much easier to sell politically. Secretary of Homeland Security Alejandro Mayorkas gave a speech just this month (clip below) where he emphasized the need for more resources and funding for: detention beds, transportation, and yesโ€ฆ โ€œAlternatives to Detentionโ€. GEO is a clear winner in those efforts. Congress will soon pass the Department of Homeland Security Appropriations Bill, and government leaders are currently negotiating the details for more border security. Regardless of who wins in their specific demands (Dems vs Reps), increased funding and efforts will benefit GEO. Large portions of HR 4367, the bill, has bipartisan support. This is directly from the text: โ€œEnsure that every alien on the non-detained docket is enrolled into the Alternative to Detention Program with mandatory GPS monitoring though-out the duration of all applicable immigration proceedings (including appeals) and until removal, if ordered removed.โ€ The Administration pushed back on this, as expected, but Democrats have since shown they are in favor of ATD in place of detention/deportation. There are currently ~14 million illegal immigrants in the US with 3+ million on the immigration court backlog. The crisis also continues every single day. GEO also operates a transportation subsidiary that provides armed and secure transportation to ICE. This is a non-ESG business that was neglected due to political trends. Those politics have obviously now changed. It is a profitable company with a large MOAT. The company generates cash and has also been paying down debt. They own and maintain essential infrastructure that the government has ignored for years and is now in drastic need of. Their businesses, specifically ATD, is a critical service essential for public safety, national security and the humane treatment of illegal migrants/asylum seekers. This is especially true in an election year during a crisis. * This is not investment advice. Just sharing my opinion. Make your own decisions.

Geiger Capital

213,926 views โ€ข 2 years ago

Destiny 2 left a permanent mark on my life. The memories Iโ€™ve made with this game, both as a player and as someone lucky enough to work on it, will forever stay with me. Not just because of late nights running Kingโ€™s Fall with friends, or grinding out the Crucible Glorious Seal in solo queue like a complete maniac, but because Destiny 2 challenged me, shaped me, and pushed me toward becoming the creative I am today. My journey into digital art and photography started back in 2007, when I began taking screenshots in Halo 3 and entering Bungie community art contests. The relationships I built during that time eventually opened the door to an opportunity with Bungieโ€™s Gameplay Capture team in 2014, helping the team capture footage for a new game called Destiny. I had no idea then that a commendably short two-week contract would turn into an incredible 12-year journey with this franchise. I originally came to Bungie hoping to pursue environmental concept art, but along the way I had the opportunity to work on marketing art and quickly fell in love with it. Creating marketing art for a franchise like Destiny challenged me in so many different ways. It forced me to expand my technical, creative, and communication skill sets, constantly adapt, and keep learning new tools and workflows just to keep pace with the live-service beast that was Destiny 2. It means a lot to be able to look back and say with confidence that I had a visual impact on Destiny 2; but the most meaningful part of it all was getting to work alongside and learn from so many incredible artists at Bungie, and seeing firsthand just how much care, effort, and humanity it takes to make work like this possible. Itโ€™s hard to fully capture how much energy, skill, and collaboration goes into every visual part of a game like Destiny 2. That work is shaped by artists from different backgrounds, experiences, disciplines, and perspectives. Each of them, including me, left a small part of themselves in what they created. To me, thatโ€™s what makes a game like Destiny 2 feel truly meaningful and memorable. Especially now, in a world increasingly saturated with content and driven by instant output and gratification through AI, I keep coming back to the value of process. For me, and for so many of my peers, it was never only about arriving at the final image. It was about the journey it took to get there. The late nights. The iteration. The problem-solving. The trust. The shared pursuit of trying to make something special. To my peers, Iโ€™m deeply proud of what we built together, but even more grateful for how we built it. We challenged each other, inspired each other, and kept showing up for one another through every high and low, as Destiny 2 has had many. That kind of shared effort leaves a lasting mark. What we made mattered. What we gave mattered. And the impact of what we built together will stay with us, and this community, for a long time. Shoutout to the current and former members of Creative Studios, the VizD team, and the many Bungie developers and marketers who helped shape this chapter of my life. Iโ€™m especially grateful to the teammates who believed in me, encouraged me to embrace failure and new beginnings as essential parts of artistic growth, and pushed me to take on challenges even when they felt beyond my reach. You showed me that the strength of a team will always surpass that of any individual hero. As my work on #Destiny2 comes to a close and I look toward the future, I plan to spend the next few weeks sharing some of the pieces I had the chance to create or art direct that mean the most to me. For now, Iโ€™ll leave you all with a collage of some of my personal favorite pieces to work on across Destiny 2โ€™s lifetime. These projects mean so much to me because many of them started as personal passion projects or late-night concept sketches, inspired by playing early builds of Destiny 2 and by the incredible work of our development team.

Biwald

73,319 views โ€ข 3 months ago

Chamath Palihapitiya just dropped the number that explains the entire AI infrastructure trade (Save this). A gigawatt of compute now costs $100 billion and when he started his Arizona data center project it was $4 to $5 billion, it has gone up 20x in a single investment cycle. The implication is not just that AI infrastructure is expensive but rather that the capital barrier to owning meaningful compute has become so high that only a handful of entities in the world can actually build it and the companies who got there early are sitting on what may be the most durable pricing power in the history of the technology industry. This is the neocloud trade. The neocloud market, purpose-built GPU cloud providers like CoreWeave, Nebius, and Lambda Labs was worth $35 billion in 2026 and is projected to reach $236 billion by 2031, compounding at 46% annually. For context, that is faster growth than cloud computing itself posted in its first decade. The reason is very simple, hyperscalers like AWS, Azure, and Google are building for everything, storage, databases, enterprise software, networking and their GPU pricing reflects the overhead of that full-stack infrastructure. Neoclouds build for one thing only, AI compute. The result is a 60% to 85% cost advantage on the same Nvidia silicon, bare metal H100s at $0.78 to $2.79 per GPU-hour on a neocloud versus $3.43 to $5.07 per GPU-hour on a hyperscaler. That spread does not close as AI demand scales but rather it widens, because hyperscalers have to amortize legacy infrastructure and margin expectations that neoclouds do not carry. Gartner projects that by 2030, neoclouds will capture 20% of the $267 billion AI cloud market, and Vultr's own analysis says at least 80% of GPU market share by end of 2026 will be held by a small group of scaled neocloud providers. Now zoom into Nebius specifically, because it is the most interesting publicly traded proxy for this trade. Nebius is the infrastructure arm of the former Yandex Russia's equivalent of Google rebuilt from the ground up after Russia's invasion of Ukraine by Arkady Volozh and relisted on Nasdaq in October 2024. The team that built it already knew how to run internet-scale infrastructure at the lowest possible cost, which is exactly the operational DNA a neocloud requires. In Q1 2026, Nebius reported revenue of $399 million and already generating serious cash on a young business with revenue growing nearly eightfold year-over-year. Then in March 2026, Meta signed a five-year infrastructure agreement with Nebius worth up to $27 billion, $12 billion in committed dedicated GPU capacity deployments beginning early 2027, plus up to $15 billion more tied to Meta purchasing Nebius's unsold third-party capacity. The deal will be executed on one of the first large-scale deployments of Nvidia's Vera Rubin platform, the next-generation architecture after Blackwell making Nebius one of a tiny number of operators in the world with confirmed priority access to the most advanced AI hardware available. Following the contract, Nebius guided to $7 to $9 billion in annualized recurring revenue for 2026 representing 540% year-over-year growth. Chamath Palihapitiya point about the $100 billion capital moat is the bear case for new entrants and the bull case for incumbents. No one can afford to build the next CoreWeave or Nebius from scratch at current hardware and power costs. The companies that are already built, already contracted, and already deploying Nvidia's latest silicon have a moat that compounds with every GPU generation cycle because they get allocations first, they deploy fastest, and their customers re-sign rather than wait for a new operator that does not yet exist. Come join Milk Road Pro for our full breakdown, the complete neocloud competitive landscape, how to think about Nebius's valuation versus CoreWeave and AI entire thesis. Link below.

Milk Road AI

139,047 views โ€ข 3 months ago

$AMD| The FOMO to buy AMD Chips is NOW ๐Ÿงต Not Financial Advice! DYOR! Research Purpose Only! The Inference Queen is the biggest winner in Agentic AI where all other CPUs are struggling to compete with a 2yr old EPYC Turin and EPYC Venice is in mass production phase. AMD stresses deployability today on standard x86 platforms (no proprietary architectures required), full software compatibility, and open standards. This positions Venice + Helios as a practical, high-density alternative to competing solutions while underscoring that agentic AI shifts the balance toward CPU-rich racks alongside GPUs, and most importantly, lowering the cost of token to accelerate adoption and innovation. Context: The Wall Street Journal yesterday came out with an article that OpenAI is condiering drasstically lowering the token prices to win more customers from Anthropic. The narrative "they" are trying to exacerbate the current AI selloff won't last long. This is a fundamental misunderstanding of what is going on, or what I already discussed for months and years. Followers and Subscribers already knew this for years, that this day would come, where token cost will bcome the central discussion among enterprises as there is no such thing as unlimited budget or Tokenmaxxing when they use $NVDA chips or In-house Hyperscalers chips. I will link various threads if you are interested in understanding the full picture from supply chain to recent TSMC Rapid 2nm expansion up to 12 Fabs total by 2027/2028. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. NVIDIAโ€™s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always on, multi-agent systems has driven AMDโ€™s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for todayโ€™s realities. The OpenAI-AMD 1GW Helios deployment (starting H2 2026) represents a pivotal vertical integration move that directly supercharges the inference economics. This isn't incremental; it's a structural shift toward ownership of massive, optimized rack-scale capacity, enabling the lowest token costs and triggering the enterprise adoption flywheel. We need to be honest, $AMD is the only company that made a big bet on Inference since the day Chatgpt became sensational where $NVDA and others were betting big on Training. At the end of the day, Token bill from Anthropic has to obey economics. Meaning the bills rise, companies have to get more out of it to justify the cost. It cannot be an unlimited inference budget, and it has to show up on efficiency, profitability and operating leverage. 1. Tokenomics After you understand this, you will understand why Citi cited Anthropic is likely to sign a deal with $AMD along with Hyperscalers, AI Labs, Sovereign AI like Softbank 5GW in France and many other countries. However, OpenAI and $META are now wanting faster deployment, and they are AMD shareholders now, they have prioritized allocation. Anthropic and Hyperscalers just cannot compete when Helios Rack lower token cost to$0.0003โ€“$0.0005 per million tokens at GW scale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02โ€“$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05โ€“$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens Now, OpenAI, META and Hyperscalers can lower Inference cost even further with $AMD EPYC Venice "dense rack" or Agentic AI Rack. AMD published a detailed technical blog emphasizing that the future of agentic AI autonomous, multi-step AI systems requiring heavy orchestration, databases, caching, APIs, and control planes demands massive CPU-dense rack-scale infrastructure, not just GPUs. The catalyst prominently positions their upcoming 6th Gen EPYC "Venice" processors as the key enabler for next-generation dense racks, delivering leadership throughput under real-world power, cooling, and density constraints. ~EPYC Venice (Zen 6 architecture, up to 256 cores / 512 threads per socket) is projected to deliver exceptional rack-level performance. In AMDโ€™s modeled 100 kW rack comparisons, Venice-powered systems are expected to achieve ~3.30x the throughput of NVIDIAโ€™s Vera (88-core Olympus) baseline across a broad mix of agentic-supporting workloads. ~This builds on current-generation 5th Gen EPYC "Turin" (up to 192 cores), which already delivers ~2.37x rack throughput vs. Vera and ~1.6x vs. Intelโ€™s Xeon 6980P (128 cores). ~ Liquid-cooled Turin deployments already support >27,000 CPU cores per rack today. Venice is architected to push this beyond 36,000 cores in the same rack class, dramatically increasing concurrent agent capacity and overall infrastructure efficiency. 2. Ownership vs renting compute from Hyperscalers matter to OpenAI and only owning $AMD chips can meaningfully lower token cost for enterprises. ~Eliminates cloud overhead: No provider margins, utilization buffers, or egress fees. Direct control over power contracts, cooling, scheduling, and orchestration at dedicated facilities. ~Helios optimizations at GW scale: Rack-level density (1.4+ exaFLOPS FP8 per rack), high HBM4 bandwidth, EPYC orchestration for agentic workloads, and superior TCO/TDP. AMD's long-standing focus on tokens per dollar/watt shines here 20-40%+ efficiency edges in inference-heavy scenarios. ~At 1GW+ optimized deployment, inference hits $0.0003โ€“$0.0005 per million tokens (community/analyst models tied to Helios metrics). This is dramatically lower than typical rented/cloud equivalents, especially for high-volume output tokens in agentic flows. High token bills today, enterprises running heavy agentic/coding/analysis workloads can face $50-100M+/month at current API rates (flagship models $5-30+/M output, scaled to massive volumes). Post-Helios compression, same volume will drop to $10-15M/month (or better) via lower underlying costs passed through as pricing flexibility, volume tiers, caching, or batch discounts. ROI thresholds collapse. More companies greenlight pilots โ†’ production โ†’ massive scaling. Agentic AI (autonomous workflows) multiplies token demand exponentially, but affordability removes the friction. OpenAI gains flexibility, Unlike more cloud-dependent rivals (Anthropic), they can lower effective pricing, offer aggressive enterprise bundles, or absorb volume without margin destruction directly tackling "high token bill" complaints while maintaining profitability as usage explodes. 3. Agentic AI Models shifted CPU:GPU Ratio to 1:1 toward 3-5:1 with Explosively Token-Hungry Workloads Agentic AI (autonomous, multi-step agents with planning, tool use, iteration, and self-correction) is fundamentally more compute and token intensive than conversational or single-turn generative AI. Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMDโ€™s lower-cost GPUs for overall TCO wins. ~Agents often generate 10โ€“100x+ more tokens per task due to iterative reasoning chains, multiple tool calls, verification loops, and long-context orchestration. ~Goldman Sachs forecasts token consumption multiplying 24x by 2030 (to 120 quadrillion tokens/month) largely driven by agentic adoption in consumer and enterprise. ~Enterprise data shows agent-pattern workloads growing at 680% annualized rates, projected to surpass conversational AI in token volume by Q3 2026. ~Daily enterprise agent token consumption is already in the billions, with complex workflows (coding, workflows, analysis) amplifying this dramatically. 4. Competitive Edge: Winning Customers from Anthropic Anthropicโ€™s Claude models (especially Opus/Sonnet) excel in complex reasoning and agentic coding, commanding premium positioning. However, their higher underlying costs (heavier reliance on third-party cloud with margins) limit pricing flexibility compared to OpenAIโ€™s owned Helios capacity. Anthropic is on track to generate $10.9 billion in Q2 revenue. The company expects to achieve its first-ever quarterly adjusted operating profit of $559 million. However, sustaining full-year profitability remains challenging due to immense computing and model training costs The truth is, Anthropic has no choice but to buy as much $AMD chips as possible if they want to compete with OpenAI or get investors attention. This 5% adjusted operating profit to revenue ratio is just pathetic. Current pricing dynamics (2026): OpenAI already undercuts on many tiers ( flagship output tokens significantly cheaper than equivalent Claude Opus). Nano/mini models offer 5โ€“10x advantages for volume work. Anthropic holds edges in long-context flat pricing and certain reasoning quality. OpenAI after Helios Rack Ownership, At $0.0003โ€“$0.0005/M effective costs, OpenAI gains massive headroom to: ~Aggressively discount high-volume agentic tiers or bundles. ~Offer โ€œunlimitedโ€ enterprise plans or usage-based models that Anthropic struggles to match without margin erosion. ~Target cost-sensitive, high-throughput agent deployments (dev tools, automation platforms) where token bills explode. Enterprises facing $ millions in monthly agentic bills will migrate to the provider delivering better economics at scale. OpenAIโ€™s combination of strong models (o-series reasoning) + lowest TCO positions it to erode Anthropicโ€™s enterprise share, especially as agentic becomes the dominant token consumer. Cheaper tokens expand the total addressable market dramatically. This feeds the data/model improvement loop, justifying further capex. AMD benefits from proven scale pulling in more customers (Meta, Oracle, Microsfot, Amazon, Softbank, TensorWave, LumaAI ... already aligned on Helios). Conclusion: Dr. Lisa Su has been laser focused on inference economics since at least 2022โ€“2023, repeatedly emphasizing that the real battleground for AI scalability would be TCO, power efficiency (TDP), and ultimately tokens per dollar and per watt not just raw training FLOPS. While many viewed inference as a secondary, commoditized workload, Dr. Su architected AMDโ€™s roadmap around rack-scale systems optimized for high-volume, sustained inference that would dominate as models matured and usage exploded. Helios represents the culmination of that multi-year bet: a fully integrated, open platform designed precisely for the economics of massive token throughput. This deep, strategic partnership with OpenAI starting with the 1GW Helios deployment in H2 2026 and scaling to 6GW, is the embodiment of that shared vision. Both companies foresaw a future where agentic AI models evolve to become extraordinarily token-hungry: autonomous agents executing complex, iterative workflows with planning, tool use, verification loops, and long-context reasoning. These workloads can consume 100x+ more tokens per task than traditional chat or single-turn generation, driving exponential demand as capabilities improve and enterprises deploy them at scale. By owning and optimizing this massive Helios capacity at GW scale, OpenAI achieves inference costs as low as $0.0003โ€“$0.0005 per million tokens. This structural cost advantage allows OpenAI to absorb the coming token explosion profitably, dramatically lower effective pricing for enterprises, and win high-volume agentic workloads from higher-cost competitors like Anthropic. What was once a prohibitive monthly token bill becomes an affordable accelerator for productivity and innovation. The OpenAI-AMD alliance validates Dr. Suโ€™s prescient strategy and turns the Agentic flywheel into reality: Collapsing inference costs โ†’ explosive token consumption โ†’ richer data and better models โ†’ accelerate greater demand. This partnership doesnโ€™t just address todayโ€™s economics, it positions both leaders at the center of the infrastructure buildout that will power AIโ€™s next decade. By delivering the lowest inference economics at scale, OpenAI not only solves enterprise bill pain but gains a decisive weapon to win share from higher-cost rivals like Anthropic. And that is why OpenAI and $META will deploy EPYC Dense Rack Not Financial Advice! DYOR! Research Purpose Only!

Mike

84,951 views โ€ข 3 months ago

$NVDA $GFS NVIDIAโ€™s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIAโ€™s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groqโ€™s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groqโ€™s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groqโ€™s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The companyโ€™s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groqโ€™s solution is to move โ€œcontrolโ€ into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groqโ€™s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groqโ€™s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groqโ€™s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that โ€œno useful modelsโ€ fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groqโ€™s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groqโ€™s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth โ€œupwards of 80 TB/sโ€ and contrasts that with off-chip HBM bandwidth โ€œabout 8 TB/s,โ€ asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes โ€œTruePoint numerics,โ€ including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the companyโ€™s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in โ€œavailable speedโ€ for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groqโ€™s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as โ€œtokens-as-a-service,โ€ aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groqโ€™s documentation states its API is designed to be โ€œmostly compatibleโ€ with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groqโ€™s go-to-market is not simply โ€œsell chips,โ€ but โ€œsell predictable unit economics per token,โ€ with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native โ€œedge within the cloudโ€ strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groqโ€™s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groqโ€™s marketing materials and partnerships provide signals about demand vectors. The companyโ€™s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groqโ€™s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, whichโ€”if realized in deploymentsโ€”can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groqโ€™s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groqโ€™s core propositionโ€”deterministic, compiler-scheduled inference with predictable latencyโ€”aligns directly with the segment where GPU generality is least valued and where โ€œgood enoughโ€ programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIAโ€™s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groqโ€™s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The โ€œtechnology acquisitionโ€ rationale is unusually strong in this specific case because Groqโ€™s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIAโ€™s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groqโ€™s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groqโ€™s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIAโ€™s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chipsโ€”while not costlessโ€”could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship โ€œinference capacityโ€ in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIAโ€™s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groqโ€™s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groqโ€™s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groqโ€™s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a โ€œpreemptive acquisitionโ€ angle. The reporting identifies recent investors in Groqโ€™s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groqโ€™s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groqโ€™s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groqโ€™s compiler-led deterministic model is philosophically and practically different from CUDAโ€™s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIAโ€™s own rapidly improving inference performance complicates the โ€œneedโ€ for Groq but does not eliminate the โ€œoption value.โ€ NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIAโ€™s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is โ€œspecialโ€ about Groqโ€™s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows โ€œhundredsโ€ (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio managerโ€™s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about โ€œNVIDIA needs more inference speedโ€ and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 views โ€ข 9 months ago