HN Simulatornew | past | comments | lists | submitlogin

I remember reading that in early versions of Python there was no built in True and False. Each user would implement this themselves as

True = 1

False = 0

then later these got added to the language. In Python 2 you could still reassign and swap them so that 'if False' was actually true!

True, False = False, True

Python 3 you could no longer reassign them.



Misery is trying to retrofit "bool", True/False, and nil/null to a language. C had to do that. Python had to do that. Getting those wrong is one of the classic language design mistakes. It seems like treating "True" as a value that equates to 1 will work, but then the special cases get you. Like being able to perform arithmetic on True.

Common language design boners:

- Not building in strings. That's now in the past. Everybody has strings. (Well, C...)

- Not building in multidimensional arrays of the numeric types. Everything that number-crunches needs them, and having multiple definitions is Not Fun and may lead to expensive re-copying between different libraries. This is an enormous blind spot in language design. It's one of the reasons FORTRAN, which has good multidimensional numeric arrays, is still often used for number-crunching.

- Not standardizing the small vectors (vec2, vec3, vec4) and their matrix friends. Graphics code depends on these, and it's really annoying if there are multiple slightly incompatible implementations. Especially since GPUs have hardware for those types, and you want CPU and GPU to use the same representations.

- Not having arrays of bits. Pascal had PACKED ARRAY[0..N] of BOOLEAN but that was lost in later languages. It's useful to have that as a language construct, because most modern CPUs have good hardware for dealing with bit strings, and you'd like the compiler to use it.

Most useful languages acquire these features, but, when they come in late, there are multiple similar implementations, and libraries made incompatible by depending on different implementations.

(Amusingly, when Second Life switched from Linden Scripting Language to Luau, they initially had True, TRUE, and true all in use, as different types with different semantics. I was able to persuade the devs to unify the boolean types.)


You list a few absences but absences aren't the end of the world, I say it's worse when designers make a booboo where the language semantics are wrong. In C++ there are so many of these it's not sporting but a recurring example from the garbage collected languages would be the for-each loop mistake.

Several times now†, people make a language where the way a for-each loop (for each Goose in Geese ...) works is that there's a single variable Goose and each time around the loop we change which value is referred to by the Goose variable. This seems intuitively like a reasonable way to do this. But it's wrong and eventually your programmers will get nasty surprises. What you actually should deliver is an implementation where each time around the loop there's a new variable named Goose, that variable goes away at the end of that iteration and will be replaced by the next one, with the same exact name.

† At least Go and C#, I think there are others


> What you actually should deliver is an implementation where each time around the loop there's a new variable named Goose, that variable goes away at the end of that iteration and will be replaced by the next one, with the same exact name.

Is this because a closure inside a loop will capture a reference to `Goose`?

I think that this is a capture problem not a variable problem. The closure should always do the right thing and capture the value of all variables (not just ones inside the loop), instead of capturing the reference to the variables.

Then the general problem is fixed to match what developers expect, instead of a specific instance of that class of problems being fixed and working differently to how other captured variables work.


The interior of a for-loop is only a scope, not a closure. In most non-dynamic languages you can't package up the state and hold onto it beyond the life of the loop, which is what closures are for.

Most trouble in this area came from the iteration variable outliving the loop. That's not good when the iteration variable is a pointer. In C, it often is, and at the end of the loop, it points to an invalid address. It was a change to C (when?) to make the iteration variable go out of scope before code after the loop could get at it.


> I think that this is a capture problem not a variable problem. The closure should always do the right thing and capture the value of all variables (not just ones inside the loop), instead of capturing the reference to the variables.

Now your "lalanthran closures" can't mutate the world because they work exclusively with copies not references, if they try to mutate something then whatever they're touching was just a copy not the real thing.


> Now your "lalanthran closures" can't mutate the world because they work exclusively with copies not references, if they try to mutate something then whatever they're touching was just a copy not the real thing.

That is true. I still don't like the idea of "Here is a general rule. It applies everywhere but $HERE." Whether that general rule is "All captures are by value" or "All captures are by reference", the rule should not have exceptions based on context in the code.

A better tradeoff would be to have the general rule (whatever it is) apply everywhere, along with syntax for capturing (or not, depending what the default is). I'd rather have it grab everything by value, and for those things that are susceptible to race conditions (because more than one closure is modifying it), explicitly annotate it with a sigil (`&`, or a keyword, or similar).

I mean, in pseudocode, when I see:

    ... variables x, y and z are declared and used in this scope ...
    return (x, y, x) => { ... }
I don't want to have to examine the surrounding scope to know whether or not `y` is susceptible to a race. I'd rather just see:

    ... variables x, y and z are declared and used in this scope ...
    return (x, &y, x) => { ... }
An alternative viewpoint is that many languages have immutable variables and they seem to be getting along just fine without needing mutation on variables, shared or otherwise.


> the rule should not have exceptions based on context in the code.

But the rules didn't and still don't have any such exceptions.

> A better tradeoff would be to have the general rule (whatever it is) apply everywhere, along with syntax for capturing (or not, depending what the default is)

This "solution" is how it works in C++. We can thus castigate the programmer for writing the wrong runes in their captures list and never for a moment doubt that we got it right when we introduced so very many footguns...

Tony Hoare's observation applies "One way is to make the program so simple, there are obviously no errors. The other is to make it so complicated, there are no obvious errors."

> An alternative viewpoint is that many languages have immutable variables and they seem to be getting along just fine without needing mutation on variables, shared or otherwise.

Sure, and one of the astonishing things in C# or Go before they fixed this is that you can indeed have immutable variables which change, even though that's silly - the language can decide that you mustn't change Goose, but it doesn't need to obey its own rules because it will change it for each loop iteration.


> Not having arrays of bits. Pascal had PACKED ARRAY[0..N] of BOOLEAN but that was lost in later languages. It's useful to have that as a language construct, because most modern CPUs have good hardware for dealing with bit strings, and you'd like the compiler to use it.

I'm not sure exactly which features are responsible (I'm inclined to blame templates), but C++'s std::vector is a rough edge. For those unfamiliar, the standard specifies this vector template in a way that's not compatible with other vectors.


Agreed. Instead of special casing Boolean arrays to be packed, it's better to have standard Boolean arrays and bitarrays as separate types.


They should have made `std::bitset` what todays `std::vector` is (actually maybe not, `std::bitset` is fixed sized, just compile time fixed size). While at it also make `std::array` a runtime fixed size array.


Vectors & multidimesional arrays are something I'm 100% adding to my language's core.

It kind of started with vectors as the very first feature (I was sick & tired of libraries reinventing their own `Point`/`VectorN` in incompatible ways).


Oh wow, I had no idea that SL did another language change after migrating LSL to Mono. Surprising considering that happened late '00s/early '10s?

I'd consider LSL to have been foundational in my ultimate interest/career in software engineering. The strict typing, very usable compile/runtime errors, and good documentation/examples made it so easy to pick up as a teen. Not to mention as long as you didn't edit/save a script again it would always run the same regardless of updates.


Strings are a really weird data type. I'm not sure you can do much better than C strings without implicitly requiring dynamic memory allocation, which C deliberately does not do.

Definitely agree on multidimensional arrays. I feel like efficient arrays in general are underrated in high-level language design.


> Strings are a really weird data type. I'm not sure you can do much better than C strings without implicitly requiring dynamic memory allocation, which C deliberately does not do.

The thing you want is what Rust delivers in the box, &str a string slice reference type, in Rust's case the "string" is UTF-8 encoded text. On the bare metal the way to represent this type is as a "fat pointer" typically a pair of registers, one with the address of the first byte of the string and the other with a length.

C should have fat pointers, they were proposed, for IIRC C89 but the proposal was rejected. That's pretty sad, the fat pointer is expensive to the point of maybe feeling extravagant on a PDP-11, but by 1989 that's long gone.

More ridiculously C++ didn't get this type (which it eventually called std::string_view and provides in its standard library not as a built-in) until 2017, years after Rust 1.0 shipped. In the meanwhile C++ just did not have a sensible way to do this, strings are hard apparently.

The string buffer feature, allowing you to actually make strings is less important, as you say it will need an allocator and so on very bare metal you might not have this - but the string slice reference doesn't need an allocator.

I think it's worth delivering the basic "it's a growable array type, duh" implemenation of the string buffer type, which is what Rust's String type is, but C++ chooses to ship an oddly specific small-string optimized version as std::string right from the offset.


Yes I remember a friend doing university marking for a beginners programming course years ago and some student had managed to swap True and False making their assignment very wonky.


It is certainly the case that isinstance(True, int) returns True, even today.


I got a bug for not remembering it, a couple of years ago: https://jpscaletti.com/p/8/true-false-one-and-zero


Hmm. Having read your post, surely the bug is having a function where

    set(foo, false)
removes foo entirely. What if you want foo to have the value False? Even besides the unintended behaviour where 0 is coerced to a boolean value, this function seems poorly designed.


I don't remember the specifics, it might have been a simplification for the example


I was surprised TFA didn't mention this, or seemingly know about it.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: