TheOtherDave comments on The Power of Reinforcement

TheOtherDave 21 Jun 2012 15:24 UTC
3 points
There’s a couple of factors here worth keeping in mind.

One is that classical conditioning continues to work, even when I’m concentrating on operant conditioning. So one result of this strategy is that my target will come to associate me with aversive stimuli, which will in turn reduce the effectiveness of my attempts at reinforcement. They will similarly associate the teaching sessions and math with those stimuli, which may be counterproductive.

Another is that a target consciously noticing my attempts at conditioning changes the whole ball game, in ways I don’t entirely understand and I’m not sure are entirely understood. Sometimes it’s a huge win. Sometimes it’s a huge lose. Staying subtle is more predictable, if I can do it, but of course it’s not always possible to avoid detection, and sometimes it’s better to admit to my attempts at conditioning than to be caught out at them. The safest move is to first establish a social context where my attempts at conditioning can be labelled “manners,” such that any attempt to call me out on them is inherently low-status, but that’s not always possible either.

When using praise signals as reinforcers for systems, like some humans, who are capable of skepticism about my motives, it helps to be seen to use expensive signals. (Attention often works well, which is one reason Internet trolls are so persistent.) Of course, that typically means I have to invest resources into my conditioning efforts.

In general, the approach I endorse is to maintain (and adjust as needed) a consistent threshold of evaluation, ignore behavior that falls below that threshold, reward behavior that clears it, and resist the temptation to go meta about the process.