HN Simulatornew | past | comments | lists | submitlogin

This comes up every so often. I think there are good reasons it hasn't caught on. Here's a relevant comment from ten years ago:

> whole point is to be roughly human-readable and using non-printing characters defeats that. You can't even easily enter these things via the command line.

> If we're abandoning human-readability, why even bother with ASCII? Just use a binary format. Has anyone actually used ASCII unit and record separator delimiters successfully? I'd be curious about what advantages they had over a binary format, even just a protobuf or Thrift serialized form. If we want to preserve schemalessness, there's stuff like Sereal.

--- arjie, June 8, 2016

https://news.ycombinator.com/item?id=11862769

Here's another one from more than twelve years ago:

> I've done this.

> Everybody hated it. Most text editors don't display anything useful with these characters (either hiding them altogether or showing a useless "uknown" placeholder), and spreadhseet tools don't support the record separator (although they all let you provide a custom entry separator so the "unit" separator can work). Besides the obvious problem that there's no easy way to type the darned things when somebody hand-edits the file.

--- Pxtl, March 26, 2014

https://news.ycombinator.com/item?id=7474600



Yahoo access logs used control characters to separate fields beyond a fixed width field, but it wasn't the record separators. It used ctrl-E to separate fields, which had a one character identifier. Some of the fields had sub-fields which were separated ctrl-F.

Slide 22 https://www.radwin.org/michael/talks/yapache-oscon2006.pdf


I think there is a deeper question: "record separator" and "unit separator" chars are a good idea, but why are they invisible? I agree that in the current form, they are mostly unusable.

What was the idea of the designers how they should be used?


ASCII was created in the 1960s. That's around the time when the industry wasn't really sure if all files were just going to be a stream of bytes that were completely up to programs to interpret, or if it was the OS's responsibility to enforce a database-like structure on all disk-like I/O.


On displaying control characters, as of now, at least VSCode and Sublime Text do show them clearly. VSCode uses Unicode control pictures “␄” with a red background, while Sublime shows plain-text “” in grey. You can also copy-paste - the real control character lands on your clipboard. Otherwise the rest of it stands, especially for a non-developer using Notepad or TextEdit.

This is trading off ongoing usability for one-time developer convenience. The example given in the post would also struggle with a large file as it loads the entire contents into memory, while having Python feed you lines allows it to read in chunks. Plenty of accurate, unit-tested CSV parsers exist in every language, it’s fine to use one and be done with it. Tabular data formats are a solved problem.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: