> Without alignment, further improvements in capability turn LLMs into wanton felony generators
Isn’t “full” alignment and guardrails a task that can never be generally achieved? One will always discover and realize the need for new guidelines and guardrails? What about vague, incomplete, inconsistent nuances and generalizations - that make this a losing battle?
I suggest ultimate safeguards need to be outside the models?
Isn’t “full” alignment and guardrails a task that can never be generally achieved? One will always discover and realize the need for new guidelines and guardrails? What about vague, incomplete, inconsistent nuances and generalizations - that make this a losing battle?
I suggest ultimate safeguards need to be outside the models?