# How to solve encoding problems on non-utf8 encoded systems

**URL:** https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034
**Category:** Support
**Created:** [January 9, 2024, 6:24am UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034 "2024-01-09T06:24:33Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![realwhzc](https://discourse.vtk.org/user_avatar/discourse.vtk.org/realwhzc/32/9186_2.png) [@realwhzc](https://discourse.vtk.org/u/realwhzc)
#### Post date: [January 9, 2024, 6:24am UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/1 "2024-01-09T06:24:34Z")

</div>

[Make SystemTools::FileExists support more languages (#19215) · Issues · VTK / VTK · GitLab (kitware.com)](https://gitlab.kitware.com/vtk/vtk/-/issues/19215)  
I submitted an issue last week, and then I learned that when reading a file `kwsys` determines the file’s existence.`kwsys` requires the string encoding to be `utf-8`.So I solved the problem by converting the string encoding to `utf8`.  
But now I’ve found that CGNS determines if a file exists via the `_access` function in io.h, encoded in `MSVC` as the local encoding `gb18030`, which leads to the fact that I can’t get both determinations to pass on this computer at the same time.  
How to solve the encoding problem.

---

<div class="post-metadata">

### Author: ![ben.boeckel](https://discourse.vtk.org/letter_avatar_proxy/v4/letter/b/ea5d25/32.png) [@ben.boeckel](https://discourse.vtk.org/u/ben.boeckel)
#### Post date: [January 10, 2024, 5:52pm UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/3 "2024-01-10T17:52:21Z")

</div>

That works for your build, but we should fix it for all builds. While we can definitely fix the vendored copy, we’ll need to get CGNS upstream to also patch it for external CGNS usages to use Windows APIs directly. We can patch ParaView’s superbuild too.

Cc: @toddy @MicK7 @mwestphal

---

<div class="post-metadata">

### Author: ![toddy](https://discourse.vtk.org/letter_avatar_proxy/v4/letter/t/edb3f5/32.png) [@toddy](https://discourse.vtk.org/u/toddy)
#### Post date: [January 11, 2024, 1:38am UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/4 "2024-01-11T01:38:03Z")

</div>

That is a terrible solution. Your VTK build has now broken the UTF-8 support internally.

@ben.boeckel It looks like `_access` should be patched to `_wacess` in combination with the use of `MultiByteToWideChar`

---

<div class="post-metadata">

### Author: ![ben.boeckel](https://discourse.vtk.org/letter_avatar_proxy/v4/letter/b/ea5d25/32.png) [@ben.boeckel](https://discourse.vtk.org/u/ben.boeckel)
#### Post date: [January 11, 2024, 2:49am UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/5 "2024-01-11T02:49:42Z")

</div>

> [@toddy](#):
>
> It looks like `_access` should be patched to `_wacess` in combination with the use of `MultiByteToWideChar`

@realwhzc, can you please try this? See `Wrapping/Tools/vtkParseSystem.c`’s `system_win32_stat` for how `_wstat` is used to see if the change works for you.

---

<div class="post-metadata">

### Author: ![realwhzc](https://discourse.vtk.org/user_avatar/discourse.vtk.org/realwhzc/32/9186_2.png) [@realwhzc](https://discourse.vtk.org/u/realwhzc)
#### Post date: [January 11, 2024, 6:10am UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/6 "2024-01-11T06:10:22Z")

</div>

Okay, I’ll try.

---

<div class="post-metadata">

### Author: ![realwhzc](https://discourse.vtk.org/user_avatar/discourse.vtk.org/realwhzc/32/9186_2.png) [@realwhzc](https://discourse.vtk.org/u/realwhzc)
#### Post date: [January 11, 2024, 8:58am UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/7 "2024-01-11T08:58:38Z")

</div>

In `cgio. c`,I copied the function `static wchar_t* system_utf8_to_wide(const char* str)` to this file,and added a line `wchar_t* wname = system_utf8_to_wide(filename);` to the function `cgio_check_file`.

Then replace `if (ACCESS (filename, 0) || cgns_stat (filename, &st) || S_IFREG != (st.st_mode & S_IFREG))` by `if (_waccess(wname, 0) || _wstat64(wname, &st))`.

After that I replaced all the `_fopen, _open` in the read related section of CGNS with `_wfopen, _wopen`, involving files `ADF_internals.c`, `ADF_interface.c`.

Now the problem is solved.

I think it’s totally a CGNS issue. However, when reading the openfoam file, an exception seems to occur as well. Although no error occurs, there is no time step in the read result, indicating that the foam file is not being read properly either.

---

<div class="post-metadata">

### Author: ![ben.boeckel](https://discourse.vtk.org/letter_avatar_proxy/v4/letter/b/ea5d25/32.png) [@ben.boeckel](https://discourse.vtk.org/u/ben.boeckel)
#### Post date: [January 11, 2024, 12:47pm UTC](https://discourse.vtk.org/t/how-to-solve-encoding-problems-on-non-utf8-encoded-systems/13034/8 "2024-01-11T12:47:28Z")

</div>

Thanks for testing. Let’s get these diffs into [CGNS itself](https://github.com/CGNS/CGNS). We can cherry-pick into VTK and apply it to ParaView’s superbuild too.
