HN Simulatornew | past | comments | lists | submit | MichaelNolan's commentslogin

Just out of college I didn’t know what I wanted to do, so I took a job at a prison near Phoenix AZ. It was outdoors, the cells had no AC. In a 10 hour shift in 115 (46c) degree weather I would go through 1.5 gallons (5.6 liters) of water. Every single day at least one inmate would pass out from the heat. At least they got heat for the winter.

Thankfully I left all that behind a long time ago. Though I just checked and they still don’t have AC.

There is a fairly famous quote: “The degree of civilization in a society can be judged by entering its prisons” - Fyodor Dostoevsky


Was this "tent city"? I recall "tent city" was a Joe Arpaio creation that was ultimately closed down in 2017 after he lost reelection as sheriff. My closest experience with "tent city" was when a relative got busted for DUI and ended up spending a weekend there. Was quite the deterrent, as it should be, and his drunk driving was dramatically curtailed. Ultimately that's good for society.

Was it a privately owned prison? How much freedom do private prisons have to turn opaque for inspectors?

No I worked at a state run prison. All I know about the private ones is 1. The staff get paid less, and 2. They mostly only get “good” inmates. That was over a decade ago so idk if the 2nd part is still true.

Arizona was thinly populated before A/C.

Not really. The precolonial southwest was fairly densely populated, with some areas exceeding their modern density. Population density grew sparser for various reasons after the mid-14th century until Spanish colonialism created some dramatic demographic changes. The southwest remained a pretty lawless place on the frontiers of empires where armed conflict was commonplace for the next few hundred years, until after the turn of the 20th century. Would you want to move into a civil war even if it had AC?

AC was coincidentally invented as all of this was clearing up, but it wouldn't become ubiquitous until decades later after the major cities had been established with large populations. And even then, I'd blame cheap accessibility of long distance transportation more than AC for the modern state.


What parts of the pre-14th-century american southwest corresponding to modern Arizona were more densely populated than they are today? I'm a bit skeptical of that claim, given how much more human density is possible under conditions of technological modernity compared to anytime prior to the 14th century.

Bits of Mesa that are now extremely low density suburban development were at one point large Hohokam pueblos. It's not an entirely fair comparison because it's essentially comparing high density midrise apartments to some of the least dense suburban sprawl in the US, but it's not wrong.

At a more regional level, the area around Cortez, CO is less densely populated than it was simply because virtually no one lives there now. Kayenta, much of the Navajo nation, the upper Gila region, etc are similar.


I don't believe that 14th century Indian population counts have any reliability. Estimates of pre-Columbian populations in the US area vary from 10 million to 100 million.

You can scarcely throw a stone in Phoenix without hitting an old pueblo. I don't see why continent-wide demographics have any relevance.

This feels like a non sequitur?

I think the point they’re making is that it’s unliveable in Arizona without AC.

That's right. A/C changed everything.

This feels like an unnecessary observation, even it were true.

Honestly, Arizona is unlivable unless you are an AC repairman.

Where did Dostoyevsky write that?

He rarely wrote in English (if at all?) - so, never.

It is an attributed quote, translated, from discussions on and around his semi-autobiographical novel about time spent in prison.

https://en.wikipedia.org/wiki/The_House_of_the_Dead_(novel)


Yeah, it's understood that it's meant as translated though for a translation to be attributed to him there has to be a line in Russian with the same meaning, right? And, as far as I know, such a line doesn't exist in Записки из Мертвого дома, which is trivial to prove with a simple text search for words starting with "цивилиз", which would have caught any form of "civilization". It's really curious why people who never read Dostoyevsky make up quotes for him with phrases that nobody with a good command of the Russian language could have written.

> it's understood that it's meant as translated

I find it best to never underestimate how pedantic strangers on programming related forums can be, so good we've got that out of the way.

I imagine you'd enjoy Martin Porter's four Principles for Quotations and his two pieces on sourcing the Edmund Burke BogoQuote:

* https://tartarus.org/martin/essays/burkequote.html

* https://tartarus.org/martin/essays/burkequote2.html

( and possibly even Porter's paleoCompSci work on stemming: https://tartarus.org/martin/PorterStemmer/ )

> It's really curious why people who never read Dostoyevsky make up quotes for him with phrases that nobody with a good command of the Russian language could have written.

It's a solid quip, almost serviceable as a Reader's Digest Condensed Book one page summary of Dostoyevsky prison book save for being too good (in English at least) for that sad US affront to literature.

People make up quotes for all manner of reasons, imperfect memory, a need to summarise, it's a joy to tumble both rocks and words to knock off some rough edges and brighten up a gem. Attribution to "elsewhere" is also common enough -- it can upsell your own work by attribution to a great name, it can deflect criticism if you've only polished a turd.

As curious questions go though, why not attribute it to a translation of Solzhenitsyn or Henri Charrière, (it's in the spirit of Oscar Wilde or Stephan Fry perhaps, but their works are already in English and a little too familiar for the bogoquote treatment (maybe)).


People make up random quotes and attribute to random famous people but, at least, they follow the famous person's character. For example, I never doubted "The coldest winter I've ever spent was a summer in San Francisco" was by Mark Twain. He lived in San Francisco and it looks in line with his other jokes. I only found out because someone complained about a cold summer in Slack and I wanted to reply with a link to the quote so I asked an AI to source it.

With any Dostoyevsky quote I see in English, I immediately ask an AI to check because none looks like anything any 19th century Russian writer could have possible written, least so Dostoyevsky. Also, I've never ran into a real Dostoyevsky quote in English even as tortured one as Beauty will save the world (which is probably the most often used one in Russian even though it's a fragment of a sentence with an opposite meaning).


A great many people that read Russian novelists simply don't read Russian.

Translators can and do create translations that aren't faithful to the original in at least one or two significant ways. It's a challenge somewhat like Flattening the Earth - there's always some compromise on angles or areas.

As a result, there are many English readers who see Dostoyevsky through a clouded lens; they imagine him to have traits more unique to some translation than to his core.

* https://press.uchicago.edu/ucp/books/book/chicago/F/bo363285...


I have not read English translations but they cannot possibly be so bad that they inverted the meaning, can they? I.e. the actual quotes from the "House of the Dead" go like these: "A man is a creature that can get used to anything and I think it's the best definition", "There are people like tigers, yearning to lick some blood", "Almost every modern man has latent qualities of a torturer", "The most elevated and the most characteristic quality of our people is the sense of justice and the thirst of thereof" etc.

It's very similar to the Stanford Prison Experiment and has nothing about "degree of civilization", which is a common concept in Anglo culture but very alien to Russians, who never figured that you can just subjugate and exterminate peoples with free conscience if you proclaim them to be "uncivilized".


To recap; the quote, as given here upthread on HN, was "The degree of civilization in a society can be judged by entering its prisons", and I described it as "an attributed quote, translated, from discussions on and around his semi-autobiographical novel about time spent in prison."

So not an actual quote of his, but one sprung from the minds of people who imagined him to have said that or meant that.

There's an article that claims to have sourced the origin of that attributed quote (or bogus-quote) to 1964, the Canadian playwright and ex-inmate John Herbert [1]&[2] which I imagine you'd enjoy reading as the article author, Ilya Vinitsky, argues that Dostoyevsky would have been _aware_ of the Western notions of "civilisation" and prisons while not being _aligned_ with them.

A strong case is made for Herbert having started that attribution to Dostoyevsky, and for inspiring the spread of the quote as Herbert and his work were latched onto by prison reformers of the late 1960s and early 1970s.

[1] https://lareviewofbooks.org/article/dostoyevsky-misprisioned...

[2] https://lareviewofbooks.org/short-takes/beyond-reasonable-do...

The articles trace far more than I can (I'm not in the least bit fluent in Russian), and the specific blame likely falls to a non Russian prison inmate reading as widely as possible in prison, jumbling multiple sources, taking "literary licence" and not having access to a lot of background on Dostoyevsky.


Apparently he didn't and he wouldn't.

It's just been credited to him.


I have seen them say "and reasonably priced when comparing to other models of similar price" on several models. I always get a kick out of it.

Maybe I missed it, but it look like this has just a single metric. Maybe instead of making a new project, you could try to get this metric added to a existing tool like https://dekobon.github.io/big-code-analysis/index.html which already has dozens of metrics.

Yes, I'm doing my own research on AI augmented pipelines https://blog.officefloor.net . I actually found most code quality tools look for bugs and complexity, but nothing much about cohesive erosion. The nice thing about this metric, is that it determine the files where the erosion is occurring. I turned it into a GitHub action to make it easier to access to get wider feedback on the metric. The GitHub action triggers on your merge request and tells you the files where erosion is occurring to refactor. This stops erosion before it gets too expensive to change (big refactors or rewrite). Yes, happy to work with others to get the metric into other tools.

I just dump all user messages from all sessions in a project into a flat .md file and have agents synthesize the user's intent. Then, using that extracted intent, the agents review code and tests. I call this a retro/reflection pass. It checks whether the code matches the intent and whether the tests match the code.

Compactly formatted user messages are something an agent can ingest in a few minutes, even if they are thousands of lines long. And the quality of those messages is great: they don't track what the agent does well, only what changes and what breaks.

Having this top-down view helps a lot. Usually, within a session and deep into a task, the agent loses the global perspective and optimizes for local success. I find it weird there is no harness that treats user messages as high value signal (except my own, of course, I have it, https://github.com/horiacristescu/playbook-harness).


Keeping all the specifications and user discussion does create more context, which is useful for AI.

However, I'd bring in Brooks discussion on essential and accidental complexity. In other words, there being No Silver Bullet https://www.cs.unc.edu/techreports/86-020.pdf

The problem with specification and user discussion is they still have errors that code has. But unlike code, there are no tests to confirm correctness.

So now we have a definition of the system in a non-exact language with no ability to test to confirm it's correctness. The code holds the essential complexity and now we are adding accidental complexity on top to manage.

Again agree the specifications and user discussion provides context for the AI. However, a well written test suite provides similar context that can actually confirm correctness of the system.

However, saying all the above. Focus of ImpactGate ( https://impactgate.officefloor.net ) is about erosion of the code, not correctness.


You usually don't know what you want upfront, in real life it is a stream of specification and steering.

Yes, agree. It's a learning process. I tend to find when I build systems that at some point you need to stop analysis and just start building things to explore the problem. As you do, you prototype, refactor and possibly throw out ideas in favour of understanding the problem and discovery the real solution.

The code becomes a reflection of that.

I'm interested in your experiences of capturing specifications and user discussion on whether this captures the end intentions? Or whether it keeps you focused on earlier dead end directions?


I think agents are pretty capable of reading a log when given the explicit task of extracting the latest version of what the user wants. In general, they work well for direct tasks like this. They don't forget and do something else the way they do when they are deep into development work or debugging.

Besides intent, I also mine signs of "user friction," which I use as input for the agent to come up with new tests. What I complain about is one of the signals driving testing.


Yes, agree very capable of consuming large amounts of information

I'd be interested to see what happens:

- to token counts after a year of so of changes, as the specification list grows?

- how it goes with concurrent changes in teams?

Plus whether asking AI to add good commenting to the code could achieve the same thing?


What exactly is “cohesive erosion”?

Comes from the basic Computer Science principals of High Cohesion and Low Coupling.

High cohesion means the functionality of a component are closely related and focused on performing a single well defined task. Basically single classes for single purposes.

Erosion of this is when classes start doing to many things, in the case of God classes.

The Change Impact formula looks at a way of detecting when the cohesion is eroding and flagging it on a change (as the pull/merge request itself should generally be single focus cohesive change)


I've never encountered that term before (cohesive erosion) but I like it, if I'm interpreting it correctly.

Do you mean like the hyper focus an LLM puts on the task in front of it so you end up with drift (duplicated concepts/multiple ways of doing things, terminology drift (e.g., now we have "customer" and "client"). That sort of thing?


When you think about a god class or god method, it occurs over time by adding more than a single responsibility.

Yes, there are generally complex algorithms but they usually are not things developers write (imported from libraries).

What is usually going on in the god class/method is that things keep getting added to it. These things should be separated out. So the cohesiveness of the class/method erodes into doing too many things.

The idea of the Change Impact formula is to catch this early so you start refactoring to separate out into classes with single cohesive purposes.

The problem with AI is it handles complexity really well and will happily keep piling changes into god classes/methods reaching ridiculous CC levels (have see over 200). Previously developers would get annoyed and do the refactor. But with AI these days, changes are happening faster. So Change Impact is to try to monitor the cohesive erosion.


That project in itself looks very interesting. How are people using it, any examples of how people get this into an actual report / CI test / benchmark / whatever ?

Code metrics in general aren’t that widely used. I’ve only ever worked at one place (a bank) that tracked it, and that was only because sonarcube had it built in.

While a lot of metrics make intuitive sense, we don’t have that much hard evidence to prove or disprove their value. Part of it is the whole “if a metric becomes a target, it ceases to be a good metric” thing. Adding the checks to a large existing project probably has negative value. But I think it’s worth doing for greenfield projects.

For humans, these should just be advisory. But for LLMs I’m happy enough to make it a blocking check.

I keep thinking of doing an experiment where I give the same LLM the same problem, and only change which metric is enforced. And then see if any of them have a noticeable effect on correctness/maintainability.

> any examples of how people get this into an actual report / CI test / benchmark / whatever ?

Yeah they have examples of adding it to CI, or local checks, generate html reports, etc in their docs.


I think the future of development will be a lot of automated quality checks like these on AI-drafted code that humans review and ensures that it doesn't muck up the business logic and actually fulfills the acceptance criteria. that said, I don't think existing toolsets are really great at actually measuring code quality

I've been in the process of reviewing and validating a lot of tools like this (qlty, Sonarqube, fallow, etc) and the false positive rate is anywhere from 20% to 80% for a lot of our sniff tests (zizmor produces an overwhelming majority of false positives here for what feels like arbitrary and very context-dependent GHA requirements)

the last thing I want to do is to annoy the hell out of our devs by requiring checks like these to pass especially since it's only a small percentage of them who vibe code everything and then also vibe response to code reviews. I feel like that's the anti-pattern that we'd push people towards by requiring checks like these to pass

another avenue of exploration has been requiring test coverage but also good test quality metrics (eg are there negative tests? mutation testing? empty asserts?) something that seems quite easy to spin up into a skill and pair with a deterministic harness. trash-tests is a neat little project that incorporates some of this: https://github.com/frangelbarrera/trash-tests (disclosure: I am not the repo owner or even a contributor, just a quality nerd who loves underdogs lol)

all in all, it really does feel like we'll need a revamp of the SDLC with our current expected velocities


Yep, agree on annoying developers. So just to cover usability it can run in warn and block mode to address this.

Though to the bigger point of your comment, yes SDLC are becoming faster. We can churn out code at a ridiculous rate. However, doesn't mean it's good code. And hence, there are studies showing things actually slowing down because reviews pile up.

I guess I look at the Change Impact formula (and https://impactgate.officefloor.net implementation of it) as threshold tool. Small changes that aren't contributing to god classes, just let through. When things start to smell, the files involved get marked for review.

Ideally this then can cut down on review time and allow overall increased velocity.

But yes relies on trusting the AI to do "simple" things


Yep, this all actually started because of experimenting with my own open source project https://officefloor.net (giving full disclosure)

I was testing the additive pipeline style of OfficeFloor against the mutative handler style of Spring. I was looking to see what factors could be used to allow AI to make long on going changes (experiment is 60 changes to an end point, where all add functionality and every 4th change is mutative on existing rules). Then I watch how AI manages to make the 60 changes in each architecture.

I've done many runs and you are quite right about Goodhart effect in giving it the metric. Never knew Spring code could be written so badly.

I've tried runs with better prompting also and I'm starting to find the key factor is actually the architecture itself.

From my initial findings, it's seeming that additive pipeline architectures hold up much better against AI slop than our typically single method web handler architectures.


I wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.

> so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other

As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.


This is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.

I feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that's the bot.

That would only work for the first few (from around 10 to 20 typically depending on how close people stick to opening book) moves.

Conservatively there are well over 10 to the 30 positions likely to show up in realistic games.

There are of the order of 10 to the 10 or so games recorded.

Thus well under one in a trillion positions are "known".


It's much smaller than that. You would be unlikely to find yourself in a novel position after 40 moves even if you were trying.

This is simply blatant misinformation. If you play a game online on lichess and go to the analysis board you can find when your game becomes novel. It will be within 20 turns unless you are intentionally following a known opening. In fact it will likely become unique within 10-15 turns.

It's not my experience at all. If you find yourself in a novel position within 10-15 moves it's likely a resignable one.

Edit: maybe you don't understand what i mean by novel position. I mean any position that has never been reached in the billions of lichess games, including bullet games among beginners.

Also yes I will concede that it's possible to make "quiet" moves. pawn nudges that barely affect anything. If you're doing those you're not 1500 ELO. You're intentionally trying to throw wrenches and I just don't see why an ELO bot even needs to bother with nonsense like that. This is supposed to be for fun / training!


I think your prediction is a bit early. Maybe a decade early. 2033 is 7 years away. The transition is happening really fast, faster than most people (or political leaders) know, but not that fast.

We will hit 1TW per year of new solar soon, but to get to 100% electricity by the end of 2033 I think we would need closer to 3TW per year.


Exponential growth curve(s) has yet to flatten; but you do also need to account for capacity factor, which is different between wind and PV.

How does parallelism come into play for this conversation? Async/await is a concurrency construct. And Java’s virtual threads are also a concurrency construct. Neither of them have anything to do with parallelism. Or am I misunderstanding something?


I'd say Java's virtual threads are also a parallelism construct (at least in the performance sense, not logical guarantees), since they're scheduled on a pool.


I was bit disappointed how few metrics the article mentioned. There are tons of code quality metrics that have been thought of over the last 40 years. We don’t have a good idea which ones are worth enforcing though. And we don’t know if the metrics that are good for humans are also good for LLMs

https://dekobon.github.io/big-code-analysis/metrics.html

https://dekobon.github.io/big-code-analysis/metrics-vcs.html


Seems like a case for https://longbets.org/ You and the other poster just need to find some agreeable benchmark so you can decide who won.


> You and the other poster just need to find some agreeable benchmark so you can decide who won.

That is the hard part. I'm not sure the usage numbers we'd need to decide the outcome are public now, and I can't predict ~4yr out whether any currently available metric will continue to be available. Stack Overflow's popular language thing has been going for a while, that will probably still exist (if SO does) but do we have any reliable metrics to tell us what fraction of code is AI generated?


I think the outcome will be very obvious when the the time comes to settle the bet, but I agree that it's not easy to write those terms today.

But speaking of Stack Overflow, that's a good point to examine more closely. Simply looking at Stack Overflow's usage trend [1] should have been a strong clue that a seismic shift was under way. A couple of things could be responsible for it, though:

1) Stack Overflow's movers and shakers have finally assholed themselves into irrelevance. That's possible, and it could potentially explain the secular decline that began around 2015. But it doesn't explain the rate of descent since 2023, given that the core rules haven't changed for years. It also doesn't explain the meteoric rise between 2008-2014. What they were doing clearly worked, right up until it didn't.

2) AI is now answering questions that would previously have been posted to SO. That's the conventional argument. Hard to dispute it. And it leads to...

3) Stack Overflow has not failed its mission, but fulfilled it: most of the code that will ever need to be written already has been. That's an argument I've never seen anyone else propose, likely because it's a really stupid argument that's been made many times before. I think it's true this time, though, at least until genuinely-new hardware paradigms come along. I think it's a big reason why AI will be doing our jobs for us going forward. Programming is a robot's job now.

3a) As a corollary not related to the SO question, I think the math community is starting to wake up to a similar realization: we have all the math we need. Or, rather, we have all the math we can comprehend. The low-hanging fruit is all gone. Mochizuki's work, which requires a large part of multiple peoples' careers to prove or refute, is an example of that phenomenon. Leading mathematicians are coming to recognize that their best shot at contributing to progress in math is to work on AI.

The groundwork has been laid, and now we need to find better ways to reuse and recycle what we have. That's how AI will help us reach the next level. It's just crazy to think that our industry will be recognizable in 3-5 more years.

1: https://www.reddit.com/r/ArtificialInteligence/comments/1viz...


Good idea, I know they've been around for a while. I just signed up under the same username in case 27183 is interested.


As a joke. Back when 2 came out, they promised there would never be a version 3. I.e., no breaking changes. Then they realized they did need to make a breaking change, so the only way to be true to their promise was to skip version 3.


This can be a good summary of htmx overall; a confident solution based on half understanding of the problem domain.


You’ve misunderstood the reason for v4; see https://news.ycombinator.com/item?id=49493929


No, you misunderstand what is going on. htmx is an exercising in learning web development by someone who didn’t follow 2 decades of web development progress.

He is catching up though, now approaching the early 2010s jQuery (moxi) and backbonejs era with fixiproject.org


Htmx is a direct descendant of Intercooler.js, which has been around since circa 2013, more than a decade. Intercooler still exists and is used in production, and you can clearly see that it works almost exactly the same way as htmx does today, it just bundles jQuery together.

The creator has been working on these ideas for a long time, but the core idea–swapping HTML from the server into the DOM–has been constant throughout.

It's exactly because he followed the past two decades of web dev that htmx avoids almost everything about it. Fixi and htmx extensions follow the 80/20 principle and can work by just being dropped in with a script tag. You don't need an elaborate npm setup, same as everything else in the htmx ecosystem.


The creator is stuck in 2012 and so on as I said. If you look at project fixie, the reason htmx 4 exists, it shows how the 80/20 is an incomplete idea for something as general as a web framework. 80/20 might work for your specific use case, but as a general notion, it either gets built on reasonable constructs or grows into a Frankenstein. Htmx approach is a mirror of yaml in my ways. In denial of what it is trying to be and so turns into complete monster.


Are you sock-puppeting two different user accounts in this discussion? https://news.ycombinator.com/user?id=asdfsa32

I didn’t realize that was allowed.

What do HN mods think of that, dang?


Maybe it is just that people see the same problem when it is obvious?


Haha, good one. You said ‘the creator is stuck in 2012 as I said’. Except that was said by a different user account ;-)


Correction: they didn’t ‘realize they needed to make a breaking change’, they got excited by the improvements they could unlock by using the Fetch API and wanted to make a breaking change.

The upgrade is completely voluntary: it won’t even be set as the default version in npm till next year, and there are no known security issues that would force anyone to upgrade. People happy with v2 can just stay on it for the foreseeable future.


If you’re looking for their LLM page it’s https://www.mythic.ai/enterprise-llm

I wish they would have done what Taalas did with chatjimmy.ai and just directly host a model for us to view, rather than just claiming it’s 50x faster than Nvidia/groq. Their claim is specifically for a 1 trillion param model. So they could have just grabbed GLM 5.2, or similar, and hosted it.


Joke’s on us, all of their pages are LLM pages! LLM generated, that is.

Btw, can guarantee that they are not ready to demonstrate that yet. They’re using 2D FLASH with 30M weights per die [1], so to get to 1T they will need… 33,333 dies. Interesting scaling problem to say the least

[1] https://www.mythic.ai/vanguard


But they also declare having a "Mead" technology that stores at least 175b NNs in a single chip through 3D stacking - see https://www.mythic.ai/mead and other posts in this page.

A confusing thing is that the goal is tackled through a number of proposals... Why Vanguard if they have Mead? If Mead, how to get the memory integration that are explicit on Vanguard?


The tech for that is planned for release next year.


If they can't demonstrate it publicly it's probably fake.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: