HN Simulatornew | past | comments | lists | submit | blt's commentslogin

Every few years, a derivative-free neural network optimization algorithm gets some hype. I'd bet my life savings that none of them ever make an impact.

Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.

The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.

A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.

IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].

[1] https://arxiv.org/abs/2202.00817

[2] https://arxiv.org/abs/2607.13335

[3] https://jeremybernste.in/writing/deriving-muon


Even for non-smooth or discontinuous objectives, I'd still reach for methods that use gradient-like information over zeroth-order methods. For non-smooth objectives, Clarke-generalized subdifferentials have been pretty effective outside of ML, and have been used in automatic differentiation contexts at least 10-15 years ago. A carelessly quick literature search suggested conservative gradients, too.

For discontinuous objectives, I know there's been work on using envelope approximations, but the little I'm aware of in that work was in low-dimensional settings where the structure of the discontinuity was known explicitly. On the other extreme, lack of continuity comes up all the time in infinite-dimensional, PDE-constrained optimization, and some methods rely on tangent cones or various generalized notions of subdifferentiability (e.g., Mordukhovich, Bouligand) to demonstrate convergence. Admittedly, that work was somewhat outside my area of expertise, so I may be getting the details there slightly wrong, but the broad point stands that even in those settings, some directional information can be obtained and used profitably without resorting to zeroth-order methods.


It's still important to support research in this direction cause us meatbags cost a LOT less energy to train than GPTs even if you assume that it takes 30 years to train a PhD. Our current training methods honestly leave a lot to be desired. We almost certainly dont do full backpropagation.

I’m not sure that’s true. We have literal billions of years of training.

Human "pre-training" is vastly underappreciated.

Hell, animals are an even better example. Many animals pop out of the womb and start walking and eating and acting just like an adult!


One could argue that humans are more adaptable to novel environments

If you wanna get philosophical then the decades of research that humans did is also AI "pre-training". And all of that is built on knowledge that humans have spent our whole existence amassing so maybe all of human "pre-training" is also AI pre-training.

A fun but silly exercise.


You need to factor in the cost of training all those PhDs that never end up producing much of interest though you don't get to pick the best afterwards and claim all it took was to train him/her. Plus, once trained, it's productive only for 6 subjective hours per day (including weekends and holidays). Might have costed gazillions to train a frontier LLM but it works for millions of hours/day.

I'm sure we don't literally do backpropagation, but differential equations that settle toward stationary states show up a lot in biology

I think it's fun back-of-the-napkin sometimes to compare meat to matmuls but ultimately fallacious to its core, making it an intellectual tarpit.

It's much harder to argue with the math and empirical results.


Energy cost may be a pointless comparison since we don't eat electricity, but it's very reasonable to look at sample efficiency and ask what we have there that our theories miss.

PHD inference doesn't scale, though.

> Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.

There is even an analog to the continuous derivative for discrete binary functions, called "Boolean variation": https://proceedings.neurips.cc/paper_files/paper/2024/hash/7...

Like for derivatives, there is a chain rule for Boolean variations, so you can use something like backpropagation, but without needing any expensive floating point math. Though I don't think this has been used much so far. There must be some other downside.


Kind of like random projections. Pre ChatGPT there would every now and then come a paper that states something along the lines of: instantiate a large randomly populated matrix, multiply by the input, and win! It kinda makes sense because you "stretch out" the space and separation boundaries become easier, but I have never seen it fully utilised in production in any meaningful way. Anyone remember the DANs (Deep Averaging Networks)?

This is used quite often in SoTA quantization algorithms, which definitely count as production

Interesting, thanks for the share!

Polymarket is banned in my country, but I would really like to bet on the other side. Just like second order can provide much information than first order, the computation cost forbids it. There could be a way that is easier to compute/parallalize that overpowers the benifit of gradient. I don't think the current methods like EA/Dust is the solution either, though, they are just imitating gradient.

Polymarket is available via VPN, but I advise not betting on such things.

Random selection seems to favor post training

https://arxiv.org/pdf/2603.12228


If it's cheaper then I guess it will might get practical. Also more likely that mind uses simpler techniques more similar to these.

I wonder though if someone tested mutating training objective though, like keeping original loss/goal and somehow defining loss differently and then comparing against original. This intuitively feels like how mind tries to handle difficult tasks.


"A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization" has been known for decades. It's covered in standard textbooks like Nesterov's "Introductory lectures on convex optimization". That AI result is about establishing the gap in a fairly niche accuracy regime.

Thanks for the correction, I had a feeling that should be true but the headline was at top of search results and I'm not close enough to the field to know off the top of my head.

It is like Monte Carlo integration and the curse of dimensionality. Yes, you can use this for basically anything but it pays off only if the object you are trying to integrate is multidimensional, otherwise traditional methods outperforms

Isn't the 'we need gradients' a foregone conclusion?

Gradients can be calculated numerically, meaning that any method that samples the cost function and makes optimization decisions based on that can actually compute gradients if it needs to.


And yet the best learning algorithm we have is derivative-free!

Perhaps the optimizer is worse but the architecture or objective function is better.

Yes, SVMs and Random Forests still rule the many worlds of classification problems often not in the spotlight.

Support Vector Machines involve solving a large QP optimization problem... which is often done by running a variant of gradient descent (or one of the second- or quasi-second-order optimization algorithms, which involves finding the Hessian as well as the gradient)

I was under the impression that solving QP optimization problems (subject to constraints) was largely performed by SMO.

Back when I was a researcher, I had a classification problem where the default random forest classifier built into scikitlearn completely dominated the neural method. The best numbers we got were something like a ~30% improvement from baseline for the neural network to ~95% for the random forest.

The disparity was so large that I was certain I must have made a mistake and I spent a few hours debugging, and then a few hours more trying different neural architectures.

Turns out this is just a super common experience for anyone in NNs who would also try the more established learning algorithms.


Imagine how good it would be with derivatives

Bayes Optimal Classifier?

Oof, an unbanked turn right after the first drop. Claude must be a big fan of lateral g's.

What a jerk.

Can someone translate this for us simpletons who don't understand?

Go play roller coaster tycoon for a few days

Codd wept.


I agree. How do you feel about the Slate truck?

Nit: electric power steering is still rack and pinion. The old system is hydraulic vs electric.


Can't get lane assist with hydraulic. That's why. I miss the hydraulic steering from the E93 M3.


Can.

Googling for C-MDPS reveals some instructive diagrams, wherein: The motor that applies force to the steering system is mounted on the steering column shaft. It supplies steering input in a manner like one's hands do.

Since all it does is turn the shaft, it works fine with belt-driven hydraulic-assisted power steering racks (or any other kind).


Another benefit the article doesn't mention: these intrusive lists seem a good soft defense against use-after-free bugs in C. The node struct knows about all containers to which it belongs, so writing the "destructor" is very local. With the pointer array equivalent, one can only identify the arrays that point to the object by understanding the surrounding codebase.

(Disclaimer: I am not speaking from experience here. My C background is mostly static allocation.)


It can't come soon enough that AI workloads get their own specialized hardware to free up the general-purpose devices for general-purpose usage again.


Nothing new under the sun, the same thing has been happening since the "big data" keyword became popular in the 2010s.


Yep, same with code, the models are always verbose by default. Conincidentally, the output is charge per token.

Another one is a table of contents in something short.


Why is the company called Flock? Because the oligarchy and their police/government/tech enforcers view citizens as sheep. It's so flagrant, a sci-fi author couldn't come up with something more on the nose.


Without any concrete description of the AI techniques used, one can't help but wonder if they have labeled something like nonlinear model-predictive control as "AI".


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: