← Back to blog
Archive · Xamarin2019-07-04 · 4 min

Realtime mobile object detector in Xamarin.Android

Archive note: This article was first published on July 4, 2019 at geeks.ms/xamarinteam (Plain Concepts Xamarin Team). It is republished here from the Wayback Machine archive 2024-05-18. Original author: Juan Antonio Cano. Code at github.com/jacano/CameraTF.

In 2019, the Plain Concepts Xamarin team contributed to TailwindTraders, Microsoft Build reference samples for Xamarin.Forms.

Our demos focused on the main Xamarin.Forms features, especially Shell.

TailwindTraders is a fictional DIY retailer. The app sells tools for work and gardening.

One part of the app was an AR experience. It read the phone’s rear camera in real time, detected a product, and showed its details with a purchase suggestion.

On Android we used the Xamarin Binding of android.hardware.camera2 for the preview, and a custom version of EmguTF to detect the objects.

We selected three objects for detection, but the demo used one: a white hardhat.

This article presents CameraTF, a Xamarin.Android sample that uses the white hardhat model from TailwindTraders.


The model

The app runs an offline model. On a phone, speed matters more than accuracy, so we chose TensorFlow Lite.

EmguTF is a C# binding for TensorFlow Lite, the C++ library for mobile and IoT. It reads a serialized model from FlatBuffer and runs the inference with the Interpreter class.

The .NET Standard project Emgu.TF.Lite is a small version of EmguTF. You give it the input tensors, and it gives you the output tensors.

For the model we used SSD MobileNet and did transfer learning over ssd_mobilenet_v1_0.75_depth_300x300_coco14_sync_2018_07_03. We trained it on Google Cloud TPUs with this pipeline config. The whole process is described in this post.

The training produced two files: hardhat_detect.tflite and hardhat_labels_list.txt.


Camera setup

The sample uses the Xamarin Binding for android.hardware.camera. That keeps the camera code short and gives us one frame at a time.

CameraTF is an Android Activity. It shows a CameraSurfaceView with the preview. We used FastAndroidCamera to get each frame.

On the devices we tested (Nokia 6.1, Google Pixel XL, LG G4), the OnPreviewFrame callback gave about 30 fps, and the preview in CameraSurfaceView stayed smooth.

CameraController sets the preview format to NV21, and the fps range and the resolution with SetPreviewFpsRange and SetPreviewSize.

The step that makes real time possible is one: convert NV21 (YUV420sp) to RGB in native code. See YuvHelper and yuv2rgb.cc.


From pixels to a detection

CameraAnalyzer runs the stages. From the RGB frame, it scales to the model input (300×300) and rotates to fix the camera orientation. SkiaSharp does the image work at native speed.

Then it fills the input tensor, calls the interpreter and reads four output tensors:

  • the bounding boxes of the detected objects,
  • the class index into hardhat_labels_list.txt,
  • the confidence of each object,
  • the number of objects.

A screen recording on a Pixel XL showed about 7 fps for the processing task inside CameraAnalyzer.

The project is open for pull requests and issues.