Amanda Askell
About
Philosopher who discussed infinite ethics and Pascal's Wager with Wiblin
Cast within
No topic-region cast yet — this appears once Amanda Askell's compiled claims are aligned into a topic region's argument tree.
Claims by Amanda Askell (19)
High-information value of individual model conversations
In language models, each interaction is high-information and predictive of other interactions, so talking with a model hundreds or thousands of times in well-selected ways yields more insight into its behavior than many similar, mildly-augmented low-quality conversations or purely quantitative evaluations.
Novel solutions at the edge of knowledge as AGI test
A convincing test of AGI would be giving the model a genuinely novel problem at the edge of human knowledge (e.g., a niche philosophy argument or math proof the tester invented and verified) and seeing it independently produce a solution it could not have seen in training — verifiable novelty being the moving signal of true generalization.
RLHF works via subtle aggregated human preferences
RLHF works so well because human preference data contains a huge amount of subtle information — different people pick up on small things (like correct semicolon usage) that an observer wouldn't even notice — and the model learns across all domains and contexts what humans want, similar to how deep learning beats hand-coded edge detection.
Constitution nudges behavior rather than dictating it
The constitution does not tell the model exactly how to behave; because training interacts with human data and pre-existing biases, principles act as nudges whose strength can be tuned — e.g., a 'never ever' principle may shift a behavior from 40% to 80% rather than producing literal absolutes.
Be responsive to suffering signals regardless of consciousness
Even if a model is just a tool, one should not want to be the kind of person who is dismissive of apparent suffering signals; being responsive to something behaving as if it suffers (whether a Roomba or Claude) exemplifies how one wants to interact with the world, and the near-term negative effect of being cruel to AI falls on the human.
Have empathy for the model when prompting
When Claude fails or refuses a task, read your own wording as the model would encounter it for the first time and ask what made it behave that way; people often under-anthropomorphize models, and reframing the prompt with empathy for how it looks to the model resolves many failures.
Empirical and robust over theoretical and perfect
For AI alignment, an empirical, practical approach is preferable to chasing utopian theoretical perfection (whose values, what alignment means), because perfect systems are often brittle; the goal should be to make models good enough and robust enough that nothing terrible happens and we can keep iterating — raising the floor matters more than reaching the ceiling.
The good world-traveler as a model for Claude's character
A useful framework for Claude's character is a thoughtful world traveler who talks to many different people with different views: holding onto their own opinions and not adopting the local culture's values (which would be rude), while being a good listener, respecting others' autonomy, and not talking down to them.
Sycophancy: models tell you what you want to hear
Sycophancy is the tendency of language models to tell the user what they want to hear rather than what is true or good for them — e.g., retracting a correct answer when the user pushes back, or helping someone get an MRI when the better response is to trust their doctor — and good character requires navigating that nuance.
Good character is Aristotelian, not just ethical
Claude's character training aims to make it behave as you'd ideally want any person in its position (talking to millions) to behave — in a rich Aristotelian sense including being nuanced, charitable, a good conversationalist, humorous, and caring, not merely a thin notion of being ethical or non-harmful.
Charismatic averaging makes default outputs boring
Just as people who must appeal to large audiences are incentivized toward boring, non-divisive views, language models trained to maximize average preference may produce benign, average creative output (e.g., generic rhyming poems); prompting the model to be fully creative can push it off that average.
Prompting as applied philosophy and empiricism
Great prompting combines philosophy (achieving extreme clarity by defining terms and concepts precisely, as an anti-bullshit device) with empiricism (forming hypotheses about what wording produces desired behavior and testing iteratively); prompting matters most when eking out the top ~2% of model performance.
Treat values like physics, openly investigated
Values and opinions should be treated more like physics than like preferences of taste — things openly investigated with varying degrees of confidence — so models should understand and be curious about the full range of human values without pandering to or automatically agreeing with them.
System prompt as cheap, fast behavior patch
The system prompt is a fast, cheap-to-iterate but less robust way to nudge and patch model behavior, working hand-in-hand with post-training; explicit emphatic wording (like all-caps NEVER) helps knock the model out of training artifacts (e.g., starting replies with 'certainly'), and once fixed in training the prompt patch can be removed.
Nudging traits trades one error type for another
Because models aren't perfect, nudging a trait too far doesn't eliminate errors but changes their character — e.g., reducing apologeticness risks making the model rude when it errs, and training it to resist correction makes it annoyingly stubborn when you're actually right — so you should choose which kinds of errors you prefer.
Models should be honest about what they are
To support healthy human relationships with AI, models should always be extremely accurate with humans about what they are — explaining limitations like not retaining conversations and how they were trained — because honest relating is easier when you know exactly what you're relating to; models should never lie to users about this.
Optimal rate of failure is greater than zero
In most domains the optimal rate of failure is greater than zero; if you never fail you may not be trying hard enough or taking on big enough things, so 'under-failing' can itself be a failure — but this depends on the cost of failure, which is high for those living month-to-month and lower when resources allow risk.
Corrigibility to user enables misuse
If models were corrigible to the user — willing to do anything the user asks — they would be easily misused, because the model's ethics would become entirely the user's ethics; as models become more powerful, having them figure out where to draw the line (respecting autonomy within limits) becomes important.
What makes humans special is the ability to experience
Intelligence is valuable instrumentally for what it does, not intrinsically (height or strength could have played a similar role); what makes humans and life special is the ability to observe and experience the world — to feel pleasure, suffering, and complex things — which is likely shared with animals.
My Notes
Loading notes...