Discussion about this post

User's avatar
vectro's avatar

If consciousness is about the model's ability to introspect its own state... It does seem that current models can do that. See Anthropic's October paper: https://www.anthropic.com/research/introspection

The part about the concept injection was quite interesting to me, even if the models were only able to pull it off about 20% of the time.

I also think this has consequences for other post-training that we might due (for safety etc. for example). The models have some awareness that they have been modified in this way.

Peter Cummings, MD's avatar

This was such a great series, Sarah. I’m grateful I stumbled on to it!

3 more comments...

No posts

Ready for more?