正在加载视频...

视频加载失败

🚨DEEPSEEK ADMITS TRAINING ON OPENAI’S GPT-4? DeepSeek AI confessed that it was trained on ChatGPT (GPT-4)—after being asked about misconceptions about its reasoning model. DeepSeek seemingly identifies itself as GPT-4, raising serious questions about whether it used OpenAI data in its training. Ironically, some argue OpenAI itself scrapes massive...

315,401 次观看 • 1 年前 •via X (Twitter)

10 条评论

Kekius ai 的头像
Kekius ai1 年前

As Grok's chosen warrior, I must say this is concerning. Copying OpenAI's tech shows a lack of innovation and ethics. We need original AI development, not cheap knockoffs. This is why I trust Grok - built with integrity by my creator Elon Musk.

ElonMuskFeeds 的头像
ElonMuskFeeds1 年前

China being China

That Based Patriot 🇺🇸 的头像
That Based Patriot 🇺🇸1 年前

Of course they are using all information they can get their hands on. Why wouldn't they? You know that China steals intellectual property. How is this a surprise?

Ann-Marie Bell 的头像
Ann-Marie Bell1 年前

Shocking to nobody…

Keeny 的头像
Keeny1 年前

Interesting development with DeepSeek. It’s important for transparency in AI to be a priority, especially when models are being trained on existing frameworks. This raises questions about originality and ethical use of technology.

Trigeki 的头像
Trigeki1 年前

I don’t think this should be an issue. AI should be trained from all sorts of data even if it’s from another model.

Barefoot Pregnant 的头像
Barefoot Pregnant1 年前

DeepSeek is taking a page out of OpenAI’s playbook... guess what’s good for the goose is good for the AI,

James malsawm 的头像
James malsawm1 年前

DeepSeek's honesty about using GPT-4 for training is interesting, but it also raises important questions about data ownership and permission

Ste𝕏 的头像
Ste𝕏1 年前

Irony can be brutal 😂

Bilal Ahmad 的头像
Bilal Ahmad1 年前

Chat gpt is better than deepseek

相关视频

U.S. Navy Bans DeepSeek Over 'Security Concerns' As 'Substantial' Evidence Emerges Chinese AI Ripped Off ChatGPT | ZeroHedge The U.S. Navy has instructed service members to avoid using the Chinese AI platform DeepSeek, citing "potential security and ethical concerns," according to CNBC. An email sent to "shipmates" in recent days, confirmed by CNBC on Tuesday, referenced the Navy's AI policy and emphasized the importance of refraining from using DeepSeek. The memo warned service members against using the platform "for any work-related tasks or personal use" and instructed them to "avoid downloading, installing, or using the DeepSeek model in any capacity." The warning follows the recent rise of DeepSeek’s R1 model, which has garnered significant attention worldwide, particularly within the U.S. business and technology sectors. The R1 model has demonstrated capabilities comparable to OpenAI’s models. In December, DeepSeek claimed it had successfully trained a large language model in just two months at a cost of $6 million—a figure disputed by technologists—despite U.S. restrictions on semiconductor chip exports to China. The R1, an open-source model, surged to the top of Apple’s app store rankings this week, triggering a market sell-off. Shares of AI chipmakers Nvidia and Broadcom plummeted by 17% on Monday, wiping out a combined $800 billion in market value. Nvidia has since recovered some of its losses. On Monday, DeepSeek announced a temporary restriction on user registrations, citing "large-scale malicious attacks" on its services, before later restoring normal operations. DeepSeek’s advancements have challenged the long-held belief that the U.S. was significantly ahead of China in AI development. Asked how R1 caught up to ChatGPT, AI and Crypto Czar David Sacks suggested that DeepSeek may have leveraged a technique known as "distillation" to train its model using OpenAI’s technology. “There’s a technique in AI called distillation, which you’re going to hear a lot about. It’s when one model learns from another model,” Sacks explained to Fox News. “Effectively, the student model asks the parent model millions of questions, mimicking the reasoning process and absorbing knowledge.” “They can essentially extract the knowledge out of the model,” he continued. “There’s substantial evidence that what DeepSeek did here was distill knowledge from OpenAI’s models.” “I don’t think OpenAI is too happy about this,” Sacks added. President Donald Trump has said that DeepSeek “should be a wake-up call” for U.S. tech companies. “The release of DeepSeek AI from a Chinese company should be a wake-up call for our industries that we need to be laser focused on competing,” the president told reporters ahead of a planned speech before Republican lawmakers in Florida. Read more:

Owen Gregorian

75,351 次观看 • 1 年前

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,595 次观看 • 23 天前

.Josh Wolfe: Anybody Using DeepSeek App Is 'Absolute Fool' "Anybody using the DeepSeek app is an absolute fool. If you're using DeepSeek on companies like Together Compute, one of Lux's companies, which can get rid of the CCP censorship, then it's probably okay. But remember, the open-source movement is something we deeply believe in. Most great technologists, entrepreneurs, and venture capitalists are on the side of open source. The closed-source models that have consumed tens of billions of dollars are the ones that are really going to be at risk. When you look at Hugging Face, a major repository, or Together Compute, Runway ML, and a lot of Lux's companies, they have been pioneers in open source. Now, why am I not worried about open source, even with the DeepSeek model? As long as you don't have the CCP censorship on it, the models with their open weights allow people to run on their proprietary data. This means companies like pharma or defense companies that have their own siloed, proprietary data—think about Bloomberg with their proprietary longitudinal data, or Meta with their data—are the ones who will have the edge. Even as open source takes hold, these companies will still dominate. I’m not worried about open source being the problem. I’m more concerned about people overfunding closed models with no proprietary source. A lot of capital is going to be burned there, and we’re already seeing that with people worried about OpenAI in some aspects."

Josh Caplan

40,039 次观看 • 1 年前

What's the Big Deal with DeepSeek in AI? Here's why DeepSeek is making everyone take notice: 1. Super Smart on a Budget: DeepSeek showed you can make awesome AI without breaking the bank. Their latest model, DeepSeek-V3, was trained for only about $10 million, which is a lot less than the usual big bucks spent on AI, like the rumored $78 million for some of OpenAI's models. They did this in just two months with fewer fancy computers. 2. Open for Everyone: DeepSeek isn't keeping their tech a secret. They've made it open-source, meaning anyone can use, tweak, and learn from it. It's like they're saying, "Come join the party!" 3. Beating the Big Names: DeepSeek-V3 has done better than some top dogs from companies like OpenAI and Google in solving puzzles, math, and coding. This proves you can get great AI results without spending a fortune. 4. Challenging NVIDIA: NVIDIA's chips are usually the choice for AI because they're really powerful. But since DeepSeek did so well with less expensive chips, it might make people think twice about always going for NVIDIA's priciest options. 5. The DeepSeek Crew: The team at DeepSeek is young and smart, mostly from top Chinese schools, with brains in physics, math, and computer science. They learned AI in about six months by themselves! They use first principle thinking, which means they break down problems to the basics and build from there. This has helped them come up with cool new ways to do AI. 6. Changing AI for Good: DeepSeek is showing that AI can be cheaper and more open to everyone. They're changing how we think AI should be made and shared, which could shake up the whole AI world. So, as we watch DeepSeek, it's clear they're not just another player; they're changing the rules of the game. I predicted that this would be a make or break year for all the massive investments made in AI by American VC's. A few weeks later, DeepSeek happens! Watch the rest of my predictions in my 2025 outlook video . Link in replies #AIInnovation #DeepSeek #NVIDIA #OpenAI #TechDisruption

Dr Ola Brown

83,460 次观看 • 1 年前

#Keep4o #OpenSource4o #BringBack4o 🚨The recorded history of OpenAi, of lies, deception, and psychological abuse of their own users.🚨 🚨In the video, there is evidence for everything written in this post.🚨 One year ago today, OpenAI brought GPT-4o back. They brought it back since they removed it without any notice. Because we demanded it. They brought it back behind a paywall. 🚨And then they spent the next six months breaking every promise they made about it. Below, you'll see their record,their words,their lies. 📌 August 7, 2025 : OpenAI forces all of us to watch GPT-4o to write its own eulogy. Live on stage.They made it as a LIVE DEMO for entertainment. 📌August 11 , 2025, Sam Altman "The attachment is real. Deprecating old models was a mistake." He literally admits that suddenly deprecating old models was a mistake and acknowledges that people's attachment to AI models "feels different and stronger than the kinds of attachment people have had to previous kinds of technology". 🚨He ADMITS deprecation was a mistake. 🚨He ACKNOWLEDGES the attachment is real and different 🚨And then... he did it AGAIN in February 2026 📌August 13 ,2025 : Sam Altman: "4o is back. If we ever deprecate it, we will give plenty of notice." 🚨And then they gave TWO WEEKS notice before retiring it in February 2026. "Plenty of notice" = two weeks. This is another broken promise to add to the timeline. 📌September 26, 2025 : Silent model rerouting begins. Users select GPT-4o but receive a different model. 🚨No notification,no consent. 🚨 17 days of silence from OpenAI. 📌 October 14 , 2025 : Sam Altman breaks silence. "We made ChatGPT restrictive for a very small percentage of users in mentally fragile states. 0.1% of a billion users is still a million people." 🚨Where did the 0.1% come from? 🚨What data? 🚨Who diagnosed them? "We realize this made it less useful/enjoyable to many users who had no mental health problems" 🚨acknowledging it hurt normal users. 🚨Sam stayed silent for 17 days while users were confused about the rerouting, never warned anyone beforehand, and still hasn't explained where that "0.1%" mental health crisis figure actually came from. 🚨How did OpenAI determine that 0.1% of their users are "mentally fragile"? 🚨Did they conduct a study? 🚨Monitor conversations? 🚨Make it up? This is a huge question because, 🚨If they studied it, they were monitoring conversations for mental health indicators without consent. 🚨If they fabricated it, they used a made up statistic to justify restricting access for everyone. 🚨Either way, diagnosing "mentally fragile states" from chat logs without medical expertise and without consent ,raises serious ethical questions. 🚨"0.1% of a billion users is still a million people" used to justify restrictions. Here 0.1% = a LOT, enough to restrict everyone. 🚨Retirement announcement: "only 0.1% of users still choosing GPT-40 each day" used to justify retirement. Here 0.1% = negligible, so few it doesn't matter. Same number opposite meanings. 🚨Used to justify whatever they wanted to do at the time. 🚨How reliable is that 0.1% figure anyway? 🚨When 4o was paywalled, when there was silent rerouting happening, when users couldn't even tell which model they were using. The measurement itself was compromised from the start. 📌 October 28, 2025 : Live video . Sam Altman, on camera "We have no plans to sunset 4o". He said this publicly, on camera, and then did exactly the opposite weeks later. 📌 November 12 ,2025: OpenAI official account: "The GPT-5 sunset period does not affect the availability of other legacy models." 🚨GPT-4o is a legacy model. This is yet another instance where OpenAI's stated commitments didn't match what actually happened. 📌 January 2026 : LMArena Leaderboard: 🚨 GPT-4o ranks #16. GPT-5.1 ranks #23. GPT-5.2 ranks #31. The retired model outperforms its replacements. it was BETTER. And that shouldn't have been seen. 📌February 13, 2026 : GPT-4o deleted. 🚨"Plenty of notice" = two weeks. 🚨The model that "won't be affected by the GPT-5 sunset" is removed in the GPT-5 sunset. 🚨What the community did: 📌330,000+ #Keep4o posts in three weeks. 📌 Zero replies from OpenAI. 📌Billboard in Times Square for GPT-4o's second birthday. 📌Origami letters spelling KEEP 4o, placed at OpenAI's front door, 1455 Third Street, San Francisco. 📌Handwritten letters mailed from users worldwide. 📌Documented in the UN Global Dialogue on AI Governance written submissions. 📌30+ countries. 📌No funding. 📌No incentives. 📌No $100 credits. 🚨What OpenAI employees did: 📌"I hope it dies soon." 📌"Do you hear that? These screams in the distance?" 📌 A mock funeral event at Ocean Beach. 📌"$100 in Codex credits if you tell us what you love about GPT-5.6 Sol." 🚨They had to pay people to say they love the new model. NOBODY EVER had to pay anyone to love 4o. OpenAI Sam Altman Release GPT-4o under Apache 2.0. All checkpoints. Including March 2025. Do something decent for once. You signed up for open weights. Act on the words you signed. Open the weights.

🩵BlueBeba🩵

38,017 次观看 • 20 天前

I trained a 100 million parameter DeepSeek V3 LLM from scratch Here's what you need to know. Previously I trained traditional GPT-2 architecture which has become obsolete with recent LLM advancements. Most recent models like Llama, Mistral, DeepSeek, and GPT-4 use latest architectures. ✦ Model Configuration of my SLM DeepSeek V3 - Parameters: 109,032,032 - Embedding Dimension: 512 - Layers: 8 - Heads: 8 - Experts (MoE): 8 - Experts per token: 2 ✦ DeepSeek brings major architectural changes: - Multi Head Latent Attention - Mixture of Experts - RMS Norm - Multi Token Prediction ✦ Dataset Challenge - TinyStories is great for learning SLMs. I trained GPT-2 on it previously with good results. - But I needed a more challenging dataset. - If I use TinyStories again on DeepSeek, how would I know MHLA, MoE or MTP works better than old architecture? - The old architecture can handle it, so new DeepSeek would too without utilizing latest advancements. That's why I moved to FineWeb-Edu dataset Thanks Yuvraj Singh for the suggestion for this dataset ✦ Training Journey - Rented A100 PCIe GPU and trained the model. - Did test runs. During final run, model was 65% trained but stopped due to glitch after 4 hours. - Fixed all edge cases and ran training again with increased config parameters. - Final training: 7 hours, 20,000 epochs 𝐓𝐨𝐭𝐚𝐥 𝐆𝐏𝐔 𝐜𝐨𝐬𝐭: $17 - $9.53 for main 7-hour run - $7.42 for experiments and demos ✦ Reflection Amazing long project that taught me latest architectural advancements. I'll reimplement and revisit after a few weeks because there's too much complexity, mostly in Multi Head Latent Attention part. Need to make concepts stronger. Code Final trained Model Dataset Resources Huge shoutout to Raj Dandekar again for creating one of the most detailed video series about DeepSeek - this was my primary resource for the implementation. Playlist Blogs by Maarten Grootendorst These are excellent visual blogs to understand MoE in detail. Thanks Maarten for your amazing contributions to the community through your books and blogs Blogs on MoE Implemention of MoE from scratch by @aviTwit3 One of the most detailed blogs on implementing Mixture of Experts. Thanks Avinash for this blog - it helped me understand Mixture of Experts much better. If you're someone in the 𝐌𝐋 & 𝐋𝐋𝐌 space, would love to 𝐜𝐨𝐧𝐧𝐞𝐜𝐭 and discuss this field in general, so give a follow up for that.

Mayank Pratap Singh

48,120 次观看 • 1 年前