perldoc perlunicode says
Character classes in regular expressions match characters instead of bytes and match against the character properties specified in the Unicode properties database.
\w
can be used to match a Japanese ideograph, for instance.
So it looks like the answer to your question is “yes”.
However, you might want to use the \p{}
construct to directly access specific Unicode character properties. You can probably use \p{L}
(or, shorter, \pL
) for letters and \pN
for numbers and feel a little more confident that you’ll get exactly what you want.