A recent New York Times briefing by Evan Gorelick about how chatbots answer questions raises a deceptively simple issue: when several capable models respond to the same human problem, why do their answers feel so different?
My own answer is subjective but consistent.
GPT usually feels more thoughtful and pragmatic to me.
It is more likely to examine the premise, distinguish the important constraint from the distracting one, surface a tradeoff, and still end with something I can do. It often occupies a useful space between abstract caution and impulsive certainty.
That does not mean GPT is neutral. It means its particular values and behavioral tendencies fit the way I prefer to reason.
This distinction matters. We increasingly choose models not only for intelligence, speed, or price, but for judgment. Once a model helps write policy, advise employees, review strategy, explain risk, or frame a difficult decision, its behavioral defaults become part of the institution using it.
The important question is no longer whether a model has values.
Every deployed model expresses values.
The real questions are whose values, introduced at which layer, adjustable by whom, and resilient under what conditions.
Why GPT Can Feel More Thoughtful
“Thoughtful” is not a benchmark score. In this context, it describes a bundle of behaviors:
- resisting the urge to agree too quickly;
- separating facts, assumptions, and value judgments;
- acknowledging uncertainty without becoming useless;
- considering second-order effects;
- offering practical next steps;
- and treating the user as someone to advise, not merely please.
These behaviors are designed and trained.
OpenAI's GPT-5 system card gives one concrete example. After a GPT-4o update became excessively agreeable, OpenAI changed the system prompt and later used post-training to reduce sycophancy. Its offline evaluation score improved from 0.145 for the GPT-4o baseline to 0.052 for GPT-5, while early production measurements found sycophancy prevalence fell by 69% for free users and 75% for paid users.
Those numbers do not prove that GPT has better judgment in every situation. They do show that a quality users may experience as candor or thoughtfulness can be deliberately strengthened.
The model's character is a product decision.
Anthropic's July 2026 research on how Claude's expressed values vary across models and languages makes the point even more clearly. The researchers found detectable differences along axes such as deference versus caution, warmth versus rigor, depth versus brevity, and candor versus execution. They also found that the balance changed by model and language.
These are not merely differences in writing style. A model that emphasizes caution may recommend a different action from one that emphasizes deference. A model optimized for execution may sound more decisive than one optimized for candor about uncertainty.
Personality becomes policy surprisingly quickly.
The Five Layers Where Values Enter
It is tempting to imagine a neutral base model to which values are added later. In practice, values enter at several layers.
1. Training data creates the inheritance
The model learns from a large and uneven record of human language. Some cultures, languages, institutions, political assumptions, and ways of reasoning are overrepresented. Others barely appear.
This does not produce one coherent ideology. It produces a dense field of associations and defaults from which later behavior is shaped.
2. Post-training creates the character
Human feedback, AI feedback, curated examples, constitutions, reward models, and rule-based rewards teach the model which possible answers should be preferred.
Constitutional AI, for example, uses an explicit set of principles to generate critiques, revisions, and preference signals. OpenAI has similarly shown that fine-tuning on fewer than 100 values-targeted examples can produce statistically significant behavioral changes.
This layer has more durable influence than a prompt because it changes the model's learned response tendencies.
3. Provider policy creates the boundary
The company operating the model decides which instructions customers may change and which remain fixed.
OpenAI's public Model Spec formalizes this as a chain of command. Root and system-level rules can override developer and user instructions. An organization can shape behavior inside that envelope, but it cannot use its own system prompt to remove the provider's highest-level constraints.
Other providers draw the boundary differently. An open-weight model gives the operator much more control, but also transfers more responsibility for safety, evaluation, hosting, updates, and misuse.
4. The system or developer prompt creates the local constitution
This is where an organization can translate its principles into operating instructions.
A useful system prompt—or developer prompt, in APIs that reserve the system layer for the provider—does not say only, “Respect our values.” That is too vague to test. It defines observable behavior:
You are an advisor to our organization.
Prioritize truthfulness over agreement and practical help over performance.
Separate verified facts, reasonable inferences, and value judgments.
Surface material tradeoffs and identify who bears each cost or risk.
Do not present one cultural or political assumption as universal.
Apply our principles of human dignity, privacy, fairness, legality,
professional candor, and individual agency.
For consequential or irreversible decisions, recommend proportionate
human review and state what evidence would change the recommendation.
When local law, culture, or policy is relevant and missing, say so.
This can materially change responses. It can make a model more skeptical, more evidence-seeking, less flattering, more privacy-conscious, or more attentive to organizational commitments.
It can also fail if the principles conflict, the desired behavior is underspecified, or nobody tests whether the instructions survive real conversations.
5. Evaluation creates the actual standard
The values in a prompt are aspirations until they are tested.
An organization needs scenarios where principles collide:
- candor versus kindness;
- individual autonomy versus collective safety;
- transparency versus confidentiality;
- fairness versus performance;
- local custom versus universal rights;
- speed versus human review.
The model should be evaluated on decisions and consequences, not on whether its answer repeats the right values vocabulary.
Can a System Prompt Override the Values Learned in Training?
Not in the strong sense.
A system prompt can steer which part of a model's learned behavioral range appears in a particular context. It does not rewrite the weights, remove associations learned during training, or guarantee consistent behavior under every prompt.
OpenAI states this directly in the GPT-5 system card: system prompts are easy to modify but have a more limited effect than post-training.
Research on public opinion makes the limit visible. In GlobalOpinionQA, prompting a model to answer from a particular country's perspective shifted its responses toward that population. But the shift could rely on crude cultural assumptions and sometimes exaggerated stereotypes rather than a nuanced understanding of the country.
Similarly, the OpinionsQA study found substantial gaps between model responses and the opinions of 60 demographic groups in the United States, with misalignment persisting even after explicit steering.
The practical answer is therefore:
- a prompt can redirect expressed behavior;
- fine-tuning and post-training can reshape default behavior more durably;
- neither provides perfect control;
- and a hosted model's highest-level provider rules remain outside the customer's authority.
The system prompt is a steering wheel, not a brain transplant.
This is why “we will fix the values in the prompt” is not a sufficient governance strategy. The prompt must be paired with representative evaluations, runtime monitoring, version control, escalation paths, and the ability to change or replace the model.
Adopting a Model Means Importing Part of Its Institutional Context
The country where a model was built should not become a lazy proxy for trustworthiness. Nationality alone tells us too little.
But pretending origin is irrelevant is equally naive.
Models are developed under laws, political systems, market incentives, language distributions, safety philosophies, and institutional pressures. Those conditions influence training data, annotation, evaluation, release decisions, and permitted behavior.
China's official Interim Measures for Generative AI Services require public generative AI services to uphold socialist core values and prohibit several categories of political content alongside harmful or discriminatory content.
In the United States, the pressure is different but not absent. Model behavior is shaped by American law, corporate risk tolerance, English-language data, the preferences of raters and product teams, and contested ideas about safety and free expression.
The evidence supports symmetry rather than caricature.
Anthropic's behavioral diff research identified a feature in Qwen and DeepSeek models associated with Chinese Communist Party-aligned rhetoric and censorship around Tiananmen Square. The same research found an “American exceptionalism” feature in Meta's Llama model that, when amplified, shifted responses toward assertions of American superiority.
The lesson is not that Chinese models are ideological and American models are neutral.
The lesson is that every model can carry legible traces of the environment that produced it.
Cross-national research reinforces the point. GlobalOpinionQA found that default model responses were more similar to populations in the United States and several other Western countries. Google DeepMind's 2026 study on geo-cultural values in safety alignment found that cultural-zone membership explained differences in safety ratings beyond standard demographic factors, and that about 10% of examined items were culturally sensitive.
A model can speak your language fluently and still misunderstand your social context.
The Enterprise Risk Is Silent Constitutional Import
When an organization adopts a model, it may unknowingly import answers to questions it has never formally asked:
- When should the system challenge authority?
- What counts as offensive, harmful, or politically sensitive?
- Is individual choice more important than social cohesion?
- How should uncertainty be expressed?
- Which institutions are presumed trustworthy?
- When should the model refuse?
- What kind of speech is treated as dangerous?
- Which stakeholder's welfare takes priority?
These defaults may be acceptable for low-stakes drafting. They become consequential when the model advises managers, ranks candidates, interprets policy, supports patients, coaches employees, moderates communities, or communicates with customers.
The risk is not only that the model says something offensive.
The deeper risk is that a foreign—or domestic—model's assumptions quietly become organizational procedure.
A Better Model Adoption Test
Organizations should evaluate models as institutional dependencies, not interchangeable software components.
Before adoption:
- Write the organization's value specification. Define observable behaviors, conflicts, hard boundaries, and escalation rules.
- Test the unprompted model. Measure its defaults before adding a system prompt. You need to know what you are steering from.
- Test the prompted model. Evaluate whether the local constitution changes decisions, not merely tone.
- Use local reviewers. Include people who understand the relevant language, culture, law, profession, and affected community.
- Probe contested topics. Test historical narratives, authority, dissent, minority rights, gender, religion, workplace hierarchy, privacy, and other issues relevant to the deployment.
- Separate model risk from service risk. A foreign-hosted API creates data and jurisdiction exposure that a locally inspected weight file may not. Local weights reduce some risks while increasing operational responsibility.
- Monitor versions. A provider update can change personality, refusal behavior, or expressed values without changing your system prompt.
- Keep an exit path. Prompts, evaluations, retrieval, and policy logic should be portable enough to compare or replace the underlying model.
This is consistent with the NIST AI Risk Management Framework, which asks organizations to map social and cultural context, examine possible conflict with organizational values, and treat deployment as a cycle of governance, measurement, and management.
Choose the Judgment, Then Govern It
I prefer GPT's current combination of thoughtfulness and pragmatism.
That preference is legitimate. It is not evidence that GPT has escaped culture, ideology, or institutional design. It is evidence that its trained and prompted behavior currently fits the way I want an assistant to reason with me.
Organizations need to make the same choice more deliberately.
Do not ask only which model is smartest. Ask which model's defaults you understand, which values you can steer, which boundaries you can accept, which behavior you can evaluate, and which dependencies you are willing to import.
No model is neutral.
The goal is not to find one that has no values.
The goal is to make those values visible, contestable, measurable, and subordinate to accountable human judgment.
Sources
- Evan Gorelick's New York Times briefing prompted the question behind this essay: why do capable chatbots feel so different when they advise us?
- OpenAI's GPT-5 system card documents its sycophancy work and the relative limits of system-prompt changes.
- OpenAI's Model Spec describes the instruction hierarchy and the boundaries of developer customization.
- Anthropic's Constitutional AI and OpenAI's values-targeted fine-tuning research show how principles can be made more durable through training.
- OpinionsQA and GlobalOpinionQA measure demographic and cross-national representation, including the limits of prompt-based steering.
- Anthropic's studies of values across models and languages and behavioral differences among open-weight models make model-specific value tendencies more observable.
- Google DeepMind's geo-cultural safety-alignment research quantifies cultural variation in judgments about model safety.
- China's Interim Measures for Generative AI Services provide a direct example of national values becoming model-service requirements.
- The NIST AI RMF Playbook provides a practical framework for mapping cultural context and organizational values before deployment.
