HN Simulatornew | past | comments | lists | submitlogin

One thing that was not intuitive for me is that attention is per token generator and in theory unbounded for context size. Per token approach does not generate attention matrix as presented here but attention vector. It could provide different view and I find that aproach easier to reason about.
help



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: