Загрузка видео...

Не удалось загрузить видео

На главную

Another piece from the Sasha Rush conversation, this time on rewards for coding RL. He said Cursor uses a mix. Some rewards look at the tool calls themselves, some only at the final output. Everything end-to-end, no process rewards guessing what happens in the middle. I agree with him...

14,826 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

I have a friend who doesn't read anything published in the past 50 years, and the more I think about it, the more I think he's onto something. The reason is that time is the best filter we have for quality. People are bad at judging quality in the moment but very good at getting rid of junk over time. — — "History is not very good at capturing all that is great in art. It is not good at that. There are many great symphonies that have been lost permanently, there are many great painters that died unknown and their paintings are gone, there's novels that have been written that no one will ever read. So history is not good at capturing all this great art. But history is very good at discarding all that is mediocre. And the amount of time that that takes, it's something like 50 years. So over the course of 50 years, what will happen is a lot of stuff that was prominent will be re-filtered and re-filtered and re-filtered, and you'll end up with a smaller group of things which have survived that test of time. So if you think about it right now, if you go back and look at the bestseller lists for 1974, 1973, there's a lot of that that would have been highly regarded at the time, which people do not read anymore for a variety of reasons, and there's some that has survived, and that's a very telling distinction. So in a world where I'm turning 60 this year, you have a limited amount of time, all four of us have active lives, we want to make sure that if we're going to sit down, we're going to read carefully, we're going to meet and we're going to discuss it in detail, we want to make sure that the work is rewarding. And the best way to ensure that is by drawing from the past." amor towles

David Perell

86,765 просмотров • 1 год назад

You know the thing with random rewards? The thing is that you never get what you want, right? And everyone is hoping for Manga Kenji Skin now... but is it really everyone? What does your heart most desire? If you look back at your life journey, are you where you wanted to be? And if not, why not? What changes would you have made in your life to be happy with yourself? Because this is what really matters. Your reaction to this reward is only a reflection of how deeply you feel inside, isn't it? Is the reward good? Is the reward bad? Isn't that only a matter of perspective? ANYWAY, enough talking, let's finally announce the reward you will receive today, after you watch this video, and after you read this text... But, what is the point of having the reward written here in this text, if it's also announced in the video AND announced in the game... isn't all of this pointless? What does it matter what door you open in the end, right? Some of them have hints, some of them don't, it's pretty much random. OR IS IT? The Mega Box door was pretty obvious, wasn't it? I guess some of them DO have hints, and some of them don't. We did say that in the video actually... the trick is to find out which ones are misleading and which ones are not. To be honest, the plain white door wasn't designed with Manga Kenji in mind... you were the ones who came up with your own theory and when it didn't land, you thought we fooled you, but in reality, you fooled yourself. But one thing I can say, we've seen a lot of crazy theories out there, and some of them are actually right, surprisingly. Maybe an accident, maybe whoever came up with that theory is a future version of me. It can happen, I read it in a book once. Or maybe not. ANYWAY, hope you have enjoyed today's reward, and let's bring the community together to vote for the best door today! #ScaryDoors

Brawl Stars

317,447 просмотров • 11 месяцев назад

In preparation for this, I checked out Dr. Jordan Cooper's (Dr Jordan B. Cooper) video on Chemnitz's examination of traditions and he made the remark that the Catholic Church contradicts the consensus of the Fathers on communion in both kinds, and that there is "no example of" communion in one kind in the Fathers. Respectfully, I think this is a messy critique which misses what is going on at the Reformation and what the precise claim of the Reformers was. I have already made a video extensively citing examples of communion in one kind starting in the 3rd century, through Tertullian, Ss. Jerome, Basil, etc, such as the widespread practice of private communion at home with the Eucharistic bread only. I will link it below. These are usually dismissed as "extraordinary instances," but this only shows that the Fathers did not actually agree with the Reformers. Remember, the Reformers (specifically the Lutheran orthodox) said that communion in one kind can NEVER happen. Johann Adam Scherzer, for example, said that if someone cannot receive wine, “[he] should be kept away from the whole, than that he should not receive the whole” (Anti-Bellarmine, disp. 11, th. 6, obj. 9). To demonstrate that the Catholic Church "contradicts" the Fathers on communion in both kinds (which is Dr. Cooper's internal critique), one needs to show that the Fathers of any age unanimously (at least morally) taught that it was a matter of faith to hold that both kinds must be received. Certainly, no consensus ever existed. On the contrary, we know they often did make the exception and allowed Christians to receive only one kind, which shows they did not agree with the Lutherans that this is a divine precept (to receive both). Dr. Cooper conflates the ordinary liturgical practice with what Catholics mean by a "consensus on the Fathers." For more, I recommend reading St. Robert Bellarmine who has a really good response to Martin Chemnitz. Here is a helpful example. It was truly a unanimous practice to preserve the Eucharist after the liturgy and treat the elements as consecrated. The Reformed (to my knowledge) unanimously deny this is the case, and the Lutherans (from my reading of Chemnitz and Gerhard) admit that Christ's presence persists but only for taking the Eucharist to the sick. The early Church affirmed Christ's presence in the Eucharist in the Catholic sense; they preserved it even for reasons which were not for the sick and nevertheless affirmed its lasting consecration (Ss. Irenaeus, Ambrose, Jerome, to name a few). The Catholic view of tradition tells us that this practice/belief (that Christ is still present post-liturgy) must be valid and true (since the universal church cannot err), but it does not establish a precept i.e. just because the whole Church often preserved the Eucharist, it does not mean we actually have to do it; it just means we can do so validly and licitly; it also means this is a certain rule which Protestants contradicted (from a Catholic POV). Even if communion in both kinds was the universal practice (which it was not), this would not suffice to say the Tridentine Church is in contradiction, since this would only tell us it is valid (duh, so does Scripture). There would need to be a unanimous testimony from the Fathers of, let's say, the fourth century, that communion in both kinds is the only way and that to oppose this is heresy; a testimony which does not exist and is even contradicted.

Kevin Fernandez

32,660 просмотров • 2 месяцев назад

THE MOST IMPORTANT Q&A OF MEDIA DAY. Mariana: You came from a very solid weekend on top of everything, but at the same time, it seems that you don't feel that the team is listening to you. Am I right? And how do you balance that? Lewis: I feel like we're going in the right direction. Rome wasn't built in one day, so it takes time to build. For me, coming into the team, I wanted to be respectful of the way they've done things in the past and just to really observe and see where our strengths and where our weaknesses are and to highlight where our weaknesses are and areas that we need to work on. But I do feel that they've been responding. I think you're starting to see, hopefully, some of the impact of the work that we're doing in the background and also into next year's car. This is a car that I've had nothing to do with in terms of developing this car over the years. Hopefully, from next year, my input goes into that car, and that will be a car that I've hopefully been a part of or will have been a part of developing. But I think we've got a really great rapport. I think we're really progressing, particularly since the summer break. I think things have started to get better, and it's all just about building trust and communication. Also, I'm coming into a team that English is not the first language, and I don't speak Italian, so it's finding a common ground. And the fact is we all want to win. We're all here to achieve the same thing, and we've got to just keep pushing. So that's why I'm trying to keep everyone motivated on difficult weekends, trying to keep everyone lifted up. But there have been many, many things we've changed this year that I suggested that they hadn't done in the past, and so they have been listening. It doesn't change straight away, just like that. It takes time to build. And as engineers, they really need proof. They need numbers. That's what they work on. So you have to sometimes push to get certain changes to be made, and then when you change it and then it works, you're like, okay. Mariana: That's what I was talking about.. Lewis: Yeah! - F1 2025 Mexico -

sim

170,303 просмотров • 11 месяцев назад

🇨🇦 Calgary Bigfoot 2014 — One of the Rare Ones 👣 I think we can all agree that most purported Bigfoot footage is either misidentified or hoaxed. That is not a dodge — it’s just the way it is. If you have watched this account for a while and wondered why I post so many clips only to tell you they are fake, that is why. The list of bad videos is huge. Real is rare. Seeing lots of fakes helps you recognize the difference. This is the original Calgary family clip from January 2014. Just a riverbank, a father, some kids, and a figure in the trees. I have known about it for years and never sat with it long enough to decide. I did that this morning. You are looking at a real Bigfoot. What sold me is the mass. This is not a thin shape sliding through brush. When it stands, you get shoulders, a thick back, and a head set low on those shoulders — the same low-set look you see on Patty and the Idaho Bigfoot. The drop from the top of the head to the hanging left hand is long. That arm length is hard to fake at this distance with a man in a suit. The hair reads as one color, reddish-brown, not clothing. Bear and moose drop out once you watch it stand and walk. The arms swing. The body stays upright. It does not bolt. It turns and leaves the way the better clips leave — calm, not panicked. The audio matters. The kids are not performing. There is no “what is that?” routine. It is “that’s not a human” and a quiet “whoa.” That sounds like people who already saw it before the camera came up, which is what the original description said. It is still a short clip at a distance. No public tracks or hair from the spot. I am not handing you a lab result. I am telling you the morphology does not leave room for the usual outs, and this is one of the few times I will say the thing on camera is the real deal. Watch the raw file. Listen to the family react. Tell me what you think.👇 #Bigfoot #Sasquatch #Calgary

Bugs Finds Bigfoot 👣🪶

3,503,188 просмотров • 14 дней назад

BREAKING.🚨 Judge Merchan has instructed the jury they do not need to have a *UNANIMOUS* verdict in order to convict former President Donald J. Trump. "One thing in particular that the judge said the jurors could do. He delivered what is being called really the pinnacle of all of this. There is no need to agree on what has occurred. They can disagree on what the crime was among the three choices." "This means they could split 4-4-4 and the judge would still treat them unanimously. What does that mean?" "Outrageous. In a normal criminal case every statutory crime has what we call elements of the offense. Like in a bank robbery case you have to rob – it has to be a financial institution, you have to show intent," said former prosecutor Andrew McCarthy. "Those are the things the jury has to agree on unanimously that they were proved beyond a reasonable doubt. Here what we’re doing is taking the element that actually makes this a felony, because remember falsification of records is normally a misdemeanor in New York. What makes it a felony is that you are concealing or committing another crime." "And here the judge is telling them they don’t have to agree about what the other crime is under circumstances where that not only is what makes this a felony, makes it a four-year potential prison penalty rather than a year or less, but it is also what gets us into the courtroom." "If this had been a misdemeanor, the time to bring this case would have lapsed in 2019. The only reason they are still able to bring this case is because it’s a felony allegedly and yet now the judge is saying you know, you don’t have to agree on what the felony is." The jury has now gone to deliberations.

Kyle Becker

5,835,329 просмотров • 2 лет назад

The most interesting part for me is where Andrej Karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > “Humans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforcement learning is terrible. It just so happens that everything that we had before is much worse.” So what do humans do instead? > “The book I’m reading is a set of prompts for me to do synthetic data generation. It's by manipulating that information that you actually gain that knowledge. We have no equivalent of that with LLMs; they don't really do that.” > “I'd love to see during pretraining some kind of a stage where the model thinks through the material and tries to reconcile it with what it already knows. There's no equivalent of any of this. This is all research.” Why can’t we just add this training to LLMs today? > “There are very subtle, hard to understand reasons why it's not trivial. If I just give synthetic generation of the model thinking about a book, you look at it and you're like, 'This looks great. Why can't I train on it?' You could try, but the model will actually get much worse if you continue trying.” > “Say we have a chapter of a book and I ask an LLM to think about it. It will give you something that looks very reasonable. But if I ask it 10 times, you'll notice that all of them are the same.” > “You're not getting the richness and the diversity and the entropy from these models as you would get from humans. How do you get synthetic data generation to work despite the collapse and while maintaining the entropy? It is a research problem.” How do humans get around model collapse? > “These analogies are surprisingly good. Humans collapse during the course of their lives. Children haven't overfit yet. They will say stuff that will shock you. Because they're not yet collapsed. But we [adults] are collapsed. We end up revisiting the same thoughts, we end up saying more and more of the same stuff, the learning rates go down, the collapse continues to get worse, and then everything deteriorates.” In fact, there’s an interesting paper arguing that dreaming evolved to assist generalization, and resist overfitting to daily learning - look up The Overfitted Brain by Erik Hoel. I asked Karpathy: Isn’t it interesting that humans learn best at a part of their lives (childhood) whose actual details they completely forget, adults still learn really well but have terrible memory about the particulars of the things they read or watch, and LLMs can memorize arbitrary details about text that no human could but are currently pretty bad at generalization? > “[Fallible human memory] is a feature, not a bug, because it forces you to only learn the generalizable components. LLMs are distracted by all the memory that they have of the pre-trained documents. That's why when I talk about the cognitive core, I actually want to remove the memory. I'd love to have them have less memory so that they have to look things up and they only maintain the algorithms for thought, and the idea of an experiment, and all this cognitive glue for acting.”

Dwarkesh Patel

1,052,518 просмотров • 11 месяцев назад