AttenzioneMNIST: A Mouse‑click Attenzione Tracciamento Set di dati Per Scritto a mano Numeri E Alfabeto Riconoscimento

Feb 22, 2024

Modelli multipli basati sull'attenzione che riconoscono oggetti attraverso una sequenza di scorci hanno riportato risultati su riconoscimento numerico scritto a mano. Tuttavia, no dati di monitoraggio dell'attenzione per numero o alfabeto riconoscimento scritti a mano sono disponibili. La disponibilità di tali dati consentirebbe di modelli basati sull'attenzione di essere valutati in confronto per le prestazioni umane. Noi raccogliamo clic del mouse attenzione tracciamento dati da 382 partecipanti provando a riconoscere numeri scritti e alfabeti (maiuscole e minuscole) da immagini tramite campionamento sequenziale. Immagini da benchmark set di dati sono presentati come stimoli. Il set di dati raccolto, chiamato AttenzioneMNIST, consiste di una sequenza di campione (mouse clic) posizioni, prEtichetta/e classe a ogni campionamento, e la durata di ogni campionamento. In media, i nostri partecipanti osservano solo 12,8% di un immagine per il riconoscimento. Proponiamo un modello di base per prevedere la posizione e la classe/i un partecipante selezionerà al prossimo campionamento. Quando esposto alle i stessi stimoli e condizioni sperimentali come i nostri partecipanti, un modello altamente citato basato sull'attenzione di rinforzo modello cadute corto di efficienza umana. 

Chinese herb cistanche

Cinese cistanche erba- Prevenire Alzheimer's Malattia prodotti

Machine learning (ML) models that recognize objects via a sequence of glimpses have gained interest in recent years due to their scalability and efficiency. Many of these models, such as 1–7, have reported experimental results on the benchmark MNIST dataset for handwritten numeral recognition. Unfortunately, no attention-tracking data for the MNIST is available. This prevents the evaluation of attention-based models in comparison to human performance. We fell into that gap by collecting a dataset from adult participants trying to recognize handwritten numerals and alphabets from images via sequential sampling. Unlike eye-movement attention tracking (emAT), a participant clicks the location in the image that he wants to see (a form of mouse-click attention tracking (mcAT)). Immediately after that, he selects the class(es) that he predicts the object might belong to based on his observations so far. Thus, at each sampling episode, our data consists of the image location selected, class label(s) predicted, and time taken since the last episode by the participant. After each image, the participant receives a reward based on his performance (accuracy and efficiency). 

Anti Alzheimer's disease

Benefici di cistanche tubulosa-Anti Alzheimer's malattia

Vantaggi di mcAT su emAT per alfabeto numerico scritto a mano 2falfabeto riconoscimento.

(1) meat contains significant intra- and inter-personal variability in fixation location, especially for static stimuli (images)8,9. So a large amount of eye fixation data is needed to reach statistically significant conclusions. mcAT is not susceptible to some of the sources of technical noise common to eye-tracking data10. (2) Eye movements can result from both voluntary and involuntary mechanisms11. To facilitate task-dependent decision-making, we present the participants with adequate time, context, and reinforcement signals, which can also be presented to an ML model. (3) The precision and accuracy of emAT data are dependent on the eye-tracker while the same of mcAT are independent of any device. (4) It is a challenge to synchronize one's eye movements with his class selection. To overcome this, in our case, the sampling location and class(es) are selected in the same episode. (5) Finally, our method allows data collection using Amazon Mechanical Turk (MTurk), as in12,13, which is cost- and time-efective, and easily reproducible.

Contributi. 

We collect a mcAT dataset, called AttentionMNIST, using MTurk from 382 participants, rewarded for accurately and efficiently recognizing handwritten numerals and alphabets (upper and lowercase) from images via sequential sampling. Images from benchmark datasets (MNIST, EMNIST) are presented as stimuli. On average, 169.1 responses per numeral/alphabet class are recorded. Using this dataset, we show the following: • On average, participants require 4.2, 4.7, and 4.9 samples to recognize a numeral, uppercase, and lowercase alphabet, which correspond to only 11.3%, 13.4%, and 13.7% of image area respectively. Classification accuracy increases with several samples. • A model, presented as the baseline, can predict the class(es) and location a participant will select at the next sampling episode with 74.4% and 67.7% accuracy respectively, both averaged over all samplings and datasets. Class prediction accuracy increases and location prediction accuracy decreases with an increase in samples. • When exposed to the same stimuli and conditions as our participants, a highly-cited reinforcement-based recurrent attention model (RAM)3 requires 3.7, 8.5, and 7.6 samples to recognize a numeral, uppercase and lowercase alphabet, which correspond to 8.9%, 21.0%, 18.7% of image area respectively. Other attention-based reinforcement models (e.g.,1,2,4,5,7,14) can be similarly evaluated in comparison to human performance. 

Cistanche supplement near me-Improve memory2

Cistanche supplemento vicino me-Miglioramento Memoria

Clicca qui per vedere  Cistanche Migliorare Memoria e Prevenire Malattia Alzheimer Prodotti

【Chiedi per più】 Email:cindy.xue@wecistanche.com /  Whats App:  0086 18599088692 /  Wechat:  18599088692

Correlati lavoro 

The temporal sequence of mouse clicks in mcAT is analogous to the eye movement scanpath10. mcAT can effectively substitute emAT as they are significantly correlated10,12,13,15–17. Different kinds of stimuli have been used in mcAT studies, such as images of animate and inanimate objects10, images of natural scenes12,13, static webpages13, search page layouts16, and two lists of alphanumeric strings for visual comparison17. However, mcAT has not been used for handwritten numeral/alphabet classification tasks or evaluation of attention-based classification models. mcAT studies have used features such as time to contact, relative fixation frequency in areas of interest (AOIs), relative proportion of subjects that clicked at least once in an AOI10, number of fixations per trial, refixation within trials, dwell times, and scanpaths17, fixation maps12,13, AOI and information flow pattern16. The sequence of time-stamped click locations and predicted class labels constitute the raw data necessary to evaluate the efficiency and accuracy of attention-based models or humans in classification tasks. Different features can be derived from this data. Our mcAT dataset, with multiple benefits over eye-tracking data, fills a crucial gap in attention-based model research in AI, ML, and other areas. Our dataset will allow attention-based models to be evaluated in comparison to human performance. Among other things, this will facilitate the development of efficient and real-time optical character recognition systems that have wide usage in practice (see for example18–20). Principles guiding visual fixations can be hypothesized and tested using our dataset. The successful principles can be carried over to develop systems for real-world visual recognition tasks where efficiency is a key concern, such as in autonomous driving. 

Dati 

I nostri dati consistono in una sequenza di T episodi per ogni partecipante. Il dato da ogni episodio consistono in (1) la posizione nella immagine cliccata dal partecipante (uno click nella immagine per episodio), (2) la/classe/e selezionata dal partecipante, e (3) il tempo impiegato dal partecipante per registrare il campione corrente (cioè il tempo tra il tempo tra il ultimo e attuale clic nella immagine ). Questa sezione espliciterà i nostri dati raccolta processo inclusi stimoli selezione, partecipanti, visivi attività, prestazioni punteggio, e dati filtraggio. 

Stimoli selezione. Stimoli sono selezionati da immagini in due set di dati benchmark (1) 

MNIST21 set di dati consiste 70,000 etichettato immagini (28×28 pixel) di 10 numeri scritti a mano {0, 1, ..., 9}. (2) 

EMNIST22 set di dati consiste di 145,600 immagini (28×28 pixel) di alfabeti inglesi scritti a mano in maiuscole e minuscole, formanti una classe bilanciata. tutte le immagini sono etichettate con una di 26 classi {a, b, ..., z}. Tuttavia, maiuscole o minuscole etichetta non è associata a qualsiasi immagine. Da ogni categoria, selezioniamo 15 numeri ben formati da MNIST e 15 alfabeti ben formati ogni da EMNIST maiuscole e EMNIST minuscole set di dati. A ben formato numerale o alfabeto è simile alla norma della sua classe. Così, noi presente stimoli da un insieme di 15(10 + 26 + 26)=930 immagini, uniche, con 15 immagini appartenenti a ciascuna delle 62 classi. Te 930 immagini ben formate sono selezionate come segue: 

Passo 1: Normalizza ogni immagine utilizzando min-max per scalare la intensità tra 0 e 1. 

Passaggio 2: Etichetta ben formata EMNIST immagini in maiuscole o minuscole. Per ogni alfabeto classe, a alfabeto ben formato da entrambe maiuscole e minuscole immagini è manualmente selezionato e etichettato. La somiglianza coseno di tutte le immagini appartenenti a quella classe con le due immagini etichettate è calcolata. Il Immagini che sono sopra la soglia di soglia coseno somiglianza (empiricamente scelto come 0.8) sono assegnate alla etichetta maiuscola o minuscola .

Passo 3: Calcolare la media delle immagini appartenenti a ogni classe. La media immagine di una classe costituisce la sua norma. Un immagine è idonea a essere un stimolo se la sua somiglianza coseno con la media immagine della sua classe è maggiore di una soglia empirica determinata (0.7 per MNIST, 0.75 per EMNIST). 

Passo 4: tra le immagini idonee, 15 immagini da ciascuna classe sono selezionate manualmente basate su come sono ben formate. Ciascuna immagine, originariamente 28×28 pixel, è ridotta a 27×25 rimuovendo i pixel vicino dei confini poiché non hanno nessuna variazione di intensità. La media di queste 15 immagini è calcolata per ciascuna delle 62 classi. Noi denotiamo queste immagini significate come I1, I2, ..., In for n classes in ogni set di dati. 

Partecipanti. 

A totale di 382 distinti individui adulti partecipanti al nostro studio. No criterio di selezione sono stati utilizzati. Un partecipante potrebbe rispondere a più immagini. Per ogni delle 62 classi, un media di 169.1 risposte sono state registrate. 

man-5989553_960_720

Benefici di cistanche tubulosa-Malattia anti Alzheimer's

Visual attività. 

The MTurk interface for our visual task is shown in Fig. 1. A canvas of size 270×250 displays a low-intensity background image at all times. The background and stimulus images are upsampled ten times to 270×250. The center of the canvas is aligned with the center of the images. Background Initially, the background is the mean of all images in the dataset from which the stimulus is drawn. After the first episode, the background is the mean of all images from the set of classes selected by the participant in the last episode. In the real world, the context for the location, size, and orientation of a numeral or alphabet is obtained from the writing in its neighborhood, which is missing here. When our experiments were conducted with a blank background, the participants often sampled locations of the image that did not contain any part of the object. This behavior was contained by presenting the mean image of the selected class(es) in a low-intensity background and reducing the size of all MNIST and EMNIST images from 28×28 pixels to 27×25. Each time the participant selects a location in the canvas by clicking on it, a 50×50 pixel patch centered at that location from the stimulus image is revealed. A patch once revealed continues to be displayed till the final episode. A participant's task consists of three steps at each episode t (t=1, ..., T): 

Passo 1: Clicca ovunque in il 270×250 tela per rivelare la patch che vuole da provare. Solo il primo clic è accettato. 

Passo 2: Riconosci il numero/alfabeto da tutti gli campioni osservati finora. il partecipante può selezionare più classi e dovrà scegliere almeno una classe da la lista di classi mostrate sotto la tela tela. 

Passo 3: Clicca "Prossimo" sulla base della schermata per procedere. A Dedurre la classe accuratamente e rapidamente, il partecipante dovrà scegliere le posizioni giudiziosamente date le sue osservazioni finché il numeroso episodio. Non c'è il limite del tempo per un episodio. Comunque, noi limitiamo il tempo totale per T episodi di una immagine a sei minuti. Scegliamo T=12 come ottimamente citato opere on basato sull'attenzione scrittura riconoscimento o generazione hanno usato meno di 12 scorci (es., RAM3 potrebbe riconoscere MNIST numeri entro 7 scorci, DISEGNARE23 potrebbe generare MNIST numeri entro 11 scorci), e umani possono riconoscere numeri scritti e alfabeti in molto meno di 12 scorci. 

Performance punteggio. A punteggio è assegnato al partecipante basato su la sua accuratezza e efficienza in termini del numero di campioni osservati. Facciamolo essere il set di classi he scelto a qualsiasi episodio t. Ten, suo punteggio at t is:

Figure 1. Our MTurk interface as seen by a participant. Te second sampling for an EMNIST uppercase alphabet is shown.

Figura 1. La nostra MTurk interfaccia come vista da un partecipante. Te secondo campionamento per un EMNIST alfabeto maiuscolo è mostrato.

image


where |.| denotes the cardinality of a set. The total score awarded in T episodes is h {{0}} T t=1 Pt. Therefore, the maximum one can score in T episodes is T if he always chooses only the correct class. The minimum one can score in T episodes is zero if he always chooses a set of classes that does not include the correct class. So, 0 Less than or equal to h Less than or equal to T. Sooner a participant selects the correct class, the higher his score will be. Thus, this scoring mechanism takes into account recognition accuracy and sampling efficiency. Trying to maximize the score by choosing only one class from the very first episode will be risky as a score of zero will be awarded if it is not the correct class, whereas a score greater than zero will be awarded if the participant chooses multiple classes (even all classes) that include the correct class. This will motivate the participant to respond based on the probable classes in his mind at any episode. The score awarded at each episode is disclosed only upon completion of T episodes to refrain from providing any hint to the participant. In MTurk, the remuneration received by a participant for an image is proportional to his total score, h. 

Dati filtraggio.

Se un partecipante' punteggio al finale (i.e. T-th) episodio per una immagine al stimolo è zero, il suo dato registrato per quella immagine è scartata. Il dato è anche scartato se un partecipante lascia il compito incompleto. Con questo criterio di selezione 2c abbiamo ottenuto risposte su 1736 stimoli da MNIST, 4431 stimoli da EMNIST maiuscolo, e 4315 stimoli da EMNIST minuscolo; che è, 169.1 risposte per classe su media.% c2�

Modelli e metodi per utilizzare dati 

In this section, we illustrate the utility of the collected data by (4.1) providing a baseline model for predicting the behavior of a participant, and (4.2) showing how an existing attention-based reinforcement model can be compared to human numeral/alphabet recognition performance. The baseline for behavior prediction. Behavior at any episode t consists of location selection and class selection. Since a sample contains different amounts of information for different observers, or even for the same observer at different times9, behavior prediction of each participant is a difficult problem. Let n be the number of classes in a dataset, ηt be the singleton set containing the true class for the stimulus image at t, ct be the set of classes and lt be the location selected by a participant at t, to be his observation at t, and 1:t denotes the sequence 1, 2, ..., t. Till any t, the observations of a participant are o1:t and the locations he selected are l1:t. We formulate the problem of a participant's behavior prediction as follows: Class prediction Estimate the probability of i∈ct (i=1, 2, ..., n) given his o1:t and l1:t, i.e. P(i ∈ ct|o1:t, l1:t). Location prediction Estimate the probability of lt+1 given his o1:t, l1:t and ct, i.e. P(lt+1|o1:t, l1:t,ct). Class prediction. To predict the class a participant will choose at episode t, we compute the probability that the image stimulus at t belongs to class I given the participant's selected locations l1:t and the corresponding observations o1:t, as follows:

image

where Ii is the mean of the stimuli images (27×25) belonging to class i, I′ is a 27×25 image containing o1:t at l1:t, · denotes scalar product, and .denotes Euclidean norm. All pixel intensities are non-negative. At any episode t, the k highest probable classes from the belief distribution P(i|o1:t, l1:t) constitute the set of classes, ˆct, predicted by our model, where k=|ct|. Te classification accuracy is measured using the Jaccard index (JI). JI measures the similarity between two sets, X and Y, as: J(X, Y) {{10}} |X ∩ Y|/|X ∪ Y|. JI is bounded between 0 and 1; if X=Y, J(X, Y)=1. At any episode t, the classification accuracy of a participant is J(ηt,ct) while that of our model is J(ηt, ˆct). Due to its denominator, JI penalizes more as the number of elements in the predicted set (ct or ˆct) that are not in ηt increases, which is a desirable property for our case. The similarity between a participant's and our model's classification is measured by J(ct, ˆct). Our model is also evaluated in terms of class selection and rejection accuracy with respect to each participant. Let st=ct − ct−1 be the set of new classes selected and rt=ct−1 − ct be the set of classes rejected by a participant at t. Similarly, ˆst=ˆct − ct−1 is the set of new classes selected, and ˆrt=ct−1 − ˆct is the set of classes rejected by our model at t. Then the model's class selection and rejection can be compared to a participant's by J(st, ˆst) when |st| > 0 and J(rt, ˆrt) when |rt| > 0, respectively. Location prediction. Hypothesis Ideally, the belief distribution over all classes should be unimodal (i.e., one peak only) and a thin Gaussian (i.e., small standard deviation) in shape indicating a participant is confident about the class (state) of the stimulus (environment). However, as evident from our data (ref. Fig. 2), a participant is often confused between multiple classes, especially during the initial few episodes. In these cases, his belief distribution has multiple peaks or is a fat Gaussian. We hypothesize a participant's goal is to converge to an unimodal and thin Gaussian, to achieve which he selectively samples locations that reduce the probability of all classes except one. This hypothesis leads to the minimization of uncertainty over the classes (environmental states) which is a well-known principle guiding action24, including eye movements25.

Figure 2. Duration and class distribution over all participants and stimuli belonging to categories '0', 'a', and 'A'.


Figura 2. Durata e classe distribuzione su tutti i partecipanti e stimoli appartenenza a categorie '0', 'a', e 'A'.

Te observations at certain locations in a stimulus image can discriminate between certain classes. Te observation at a location l might indicate that the numeral/alphabet belongs to class I and not to class j. Such locations are more salient than others in achieving a participant's goal. To sample such locations, a saliency map, Dij, is computed such that if l is salient, the observation at l is evidence to increase the probability of class I and decrease that of j. Mathematically, Dij = N (., σ ) ∗ g(.), where ∗ is the convolution operator, g(.) is a saliency scoring function, and N (., σ ) is a 5×5 Gaussian kernel with standard deviation σ = 6 to smooth the saliency scores. We denote the set of all saliency maps as D = {Dij: i, j ∈ {1, 2, ..., n}, i �= j}. A location l in a stimulus image is salient for class i with respect to class j if Dij(l)>θ, dove la soglia θ=0.5 × max(D) è una quantità scalare determinata empiricamente.

We consider two asymmetric metrics, Kullback-Leibler (KL) divergence and difference, as candidates for the function g. KL divergence Given two normalized mean images, Ii and Ij, the KL divergence KL(Ii, Ij) measures the loss of information when Ij is used to approximate Ii. This is calculated for each pixel k as26: KL(Ii,k, Ij,k)=Ii,k log δ + Ii,k Ij,k+δ, where Ij,k is the intensity of the kth pixel of Ij, and δ is a regularization constant. When Ii,k=Ij,k, KL(Ii,k,Ij,k) → 0. Difference Given two normalized mean images, Ii and Ij, the difference for each pixel k is Diff (Ii,k, Ij,k)=Ii,k − Ij,k. When Ii,k=Ij,k, Diff (Ii,k, Ij,k)=0. A participant is uncertain regarding the set of classes, ct, he selected at the current episode. Hence, for location prediction, we consider only those saliency maps in D that involve the classes in ct. A location is predicted if it is salient based on these saliency maps and was never selected by the participant. Tus, given o1:t, l1:t and ct, the location lt+1 is predicted as follows:

image

where Ŵ is the set of 3-tuples containing the predicted location ˆl, the class it is salient for (i), and with respect to which class (j). Te location is predicted correctly if there exists a �ˆl, i, j� ∈ Ŵ such that �ˆl − lt+1� < ǫ, I ∈ ct+1 and j /∈ ct+1, where ǫ is the maximum Euclidean distance between the center pixel and any pixel in an observation patch. Te pseudo code for location prediction is shown in Algorithm 1. A detailed explanation of the pseudo-code is included in Section S1 of the supplemental material. (Te probability distribution, P(lt+1|o1:t, l1:t,ct), may be computed by assuming the saliency score of locations not in Ŵ to be zero, and then normalizing the saliency score of all locations to sum to unity. However, this probability has not been used, as Eq. (3) is sufficient for the purposes of this paper.)

image

Valutazione di attenzione‑modelli basati. 

As a representative of attention-based models, we consider the highly-cited recurrent attention model (RAM)3 that reports experimental results on the MNIST dataset. Tis reinforcement model sequentially samples an image and decides where to sample next at each sampling instant, making it appropriate for evaluation using the collected data. 

RAM 

classifiche immagini utilizzando una sequenza di scorci. Te posizione successiva è scelta stocasticamente da un distribution parametrizzata da una posizione rete. Te modello è addestrato end-to-end massimizzando il seguente obiettivo3 :

image


dove M è il numero di episodi, T è il numero di osservazioni, xi 1:t è la sequenza di interazione ottenuta eseguendo il corrente agente finché I episodi, ui t è la azione corrente, θ è il set di parametri addestrabili, Ri t è la ricompensa cumulativa, bt è una baseline, e π(ui t|xi 1:t; % ce� ) è la politica. RAM's comportamento può essere confrontato con i partecipanti' confrontando le mappe di fissazione ottenute dalla sequenza di posizioni previste da RAM e quelle scelte da i partecipanti. Un fxazione mappa è calcolata da assegnando ogni posizione un valore uguale alla frequenza della sua selezione, e poi normalizzando quelli valori per creare una distribuzione su tutte le posizioni.

Metriche per confronto fissazione mappe. Per metriche confronto due fissazioni mappe, P e Q, noi da vicino seguiamo 26. Usiamo tre metriche basate sulla distribuzione: KL divergenza (KL), Pearson coefficiente di correlazione (CC), e Somiglianza (SIM), per confrontare la distribuzione di campionamento posizioni da un modello con quella dal numero partecipanti come registrato nei dati raccolti. 

KL (definito precedente) è molto sensibile a valori zero. 

CC può valutare la relazione lineare tra due mappe come26: CC(P, Q)=σ (P, Q) σ (P)σ (Q), dove σ è la varianza o covarianza. Poiché CC è simmetrica, non neces a deduzione se differenze tra fissazione mappe sono dovute a falsi positivi o falsi negativi. 

SIM is measured as 26: SIM(P, Q)=k min(Pk, Qk), where k Pk=k Qk=1. Like CC, SIM is symmetric and inherits the same drawback. Also, SIM is very sensitive to missing values and penalizes predictions that fail to account for the ground truth density. 

Ricerca Umana e Animale. 

Il Consiglio Istituzionale della Recensione Istituzionale della Università di Memphis ha determinato che questo studio non incontra il Ufficio di Soggetti Umani Ricerca Protezioni definizione di soggetti umani ricerca e 45 CFR parte 46 non applica. Quindi, questo studio non richiede IRB approvazione o revisione. 

Sperimentale risultati dati analisi. 

The collected data can be visualized in terms of the sequence of distribution of selected locations (Fig. 3), selected classes (Fig. 2), and duration between consecutive episodes (Fig. 2). These distributions are very similar for the three datasets. For any numeral or alphabet, the distribution of selected locations after the final episode resembles the distribution of pixel intensities of its class from the dataset. However, the sequence of locations selected is stochastic in nature. The class distribution indicates confusion between categories with similar structures in the initial few episodes when the participants choose multiple classes. This confusion is reduced with more sampling. There is a significant positive correlation between the degree of confusion (# selected classes/total # classes) and sampling duration (see Fig. 4). If the number of selected classes is high (low), the duration between consecutive episodes is high (low). The CC of the sequence of locations selected by a participant for a class is not significant (Table 1). This is expected due to inter-subject variability in sampling static images. The average number of samplings required by a participant to accurately predict a class is quite low. On average, it takes 4.2, 4.7, and 4.9 samples corresponding to 36, 44.1, and 48.1 seconds to accurately classify MNIST, EMNIST uppercase and lowercase images respectively. The participants on average viewed only 11.3%, 13.4%, and 13.7% of the image area for classifying a numeral, uppercase, and lowercase alphabet image accurately (see Fig. S2 in the supplemental material). These results highlight the efficiency of the human visual reasoning system, albeit at a lower resolution than eye-tracking data but with less noise and variability. These empirical results may be useful for designing attention-based models for real-world applications. Behavior prediction. In this section, the performance of our baseline model is evaluated in terms of how accurately it can predict each participant's location and class selection. Since our experimental results using the two saliency scoring functions, KL divergence, and difference, are quite similar, results are reported using difference only, unless otherwise stated. Class prediction. The class prediction and its accuracy evaluation methods are described in the "Class prediction" section. The class prediction accuracy, shown in Fig. 5, is computed over all classes for all samplings. The mean class prediction accuracy over all samplings and datasets is 74.4% (std. dev. 26.5). Figures 5a, and b show that the set of classes selected by the participants and by our baseline model (Eq. 2) is quite inaccurate at the initial episodes and improves with increase in samples. Figure 5c shows that, during the initial episodes, these two sets, ct, and ˆct, are quite dissimilar; similarity increases with an increase in samples. The same applies to new class selections (ref. Fig. 5f). However, class rejections are similar at the initial episodes; similarity increases further with more samples (ref. Fig. 5e). Since J(st, ˆst)=|(ct ∩ ˆct) − ct−1| |(ct ∪ ˆct) − ct−1| and J(rt, ˆrt)=|ct−1 − (ct ∪ ˆct)| |ct−1 − (ct ∩ ˆct)|, it can be inferred from Fig. 5e, f that at the initial episodes, the intersection between ct−1 and ct ∪ ˆct is small, indicating that initially the participants and our baseline model make many changes in their class selection between consecutive episodes. Therefore, initially, the class selection process is highly stochastic. While there are some dissimilarities between the participants' and our model's class prediction during the initial episodes, the behaviors become increasingly similar with more samples. During the first few (typically 4 to 7) episodes, highly salient parts of a stimulus are revealed. This helps to select only the correct class in the later samplings, which increases the prediction accuracy. Since there are many classes whose mean templates match the observed parts of the stimulus during the initial few episodes, the class selection process is significantly more stochastic, leading to low classification accuracy from the participants as well as our model.

Figure 3. Distribution of sampling locations over all participants for each numeral/alphabet class and each sampling episode. Each row corresponds to a class, each column corresponds to a sampling episode which increases from left to right.


Figure 3. Distribution of sampling locations over all participants for each numeral/alphabet class and each sampling episode. Each row corresponds to a class, each column corresponds to a sampling episode which increases from left to right.

Location prediction. Our baseline model's (Eq. 3) location prediction accuracy, averaged over all samplings and datasets, is 67.7% (std. dev. 14.1) (ref. Fig. 5d). The trend of this prediction accuracy is opposite to that of class prediction accuracy. However, the explanation remains the same. Location prediction accuracy is high during the initial samplings because during these episodes, the highly salient locations are selected, leaving the less salient locations to be selected in the later episodes. Since there are many locations with low saliency, their selection process is highly stochastic and hence difficult to predict, leading to a decrease in prediction accuracy with an increase in samplings. The decreasing trend is unique for each dataset (ref. Fig. 5d) as the number of classes and the number of highly salient locations useful for discrimination vary between datasets. The lower the number of classes and highly salient discriminative locations, the faster will be the decrease in location prediction accuracy with an increase in samplings.

imageFigure 4. (Lef) Errorbar plot of time diference (seconds) between consecutive samples averaged over all classes. Tat is, value shown at sampling episode t is the time elapsed between a participant's clicks in image at t − 1 and t. (Right) Errorbar plot of confusion averaged over all classes at each episode. Errorbars indicate std. dev.

Figura 4. (Lef) Barra di errore tracciato di tempo differenza (secondi) tra campioni consecutivi media su tutte classi. Tat è, valore mostrato a campionamento episodio t è il tempo tra tra un partecipante's clic in immagine a t − 1 e t. (Destra) Barra di errore trama di confusione media over tutte classi a ogni episodio. Barre di errore indicare std. dev.

Figure 5. Evaluation of our baseline model (ref.

Figura 5. Valutazione del nostro modello di base (rif. "Baseline per comportamento previsione" Sezione). (a) Classificazione accuratezza (acc.) degli partecipanti e (b) che del nostro modello di base con etichette effettive come verità di base. (c) Classificazione somiglianza (J(ct, ˆct)), (d) posizione previsione accuratezza, (e) classe rifiuto accuratezza e (f) classe selezione accuratezza del nostro modello di base con partecipanti' dati come verità di base. Vedere % 22Comportamento previsione" sezione per dettagli.

Table 1. Average Pearson correlation coefficient (corr.) for fxation sequences for the same class. For any fixation, distance is Euclidean and direction is measured as the polar angle with respect to the center of stimuli as the origin. Std. dev. are included in parenthesis.


Tabella 1. Media Pearson correlazione coefficiente (corr.) per sequenze fxation perla stessa classe. Per qualsiasi fissazione, distanza è e e direzione è misurata come il numero angolo polare con rispetto del centro degli stimoli come la origine. Std. dev. sono inclusi in parentesi.

Valutazione di RAM. 

For each class and sampling, the fixation maps from RAM (we used the RAM implementation from github.com/hehefan/Recurrent-Attention-Model) and the collected data for the same stimuli presented in MTurk are compared. For a fair comparison with the participants, in RAM we fixed the sequence length at T=12, the first sampling location at the image center, the input observation to a 5×5 patch with the selected location as its center, and modified the reward function by Eq. (1). Te cumulative reward, Rt in Eq. (4,) is replaced by the cumulative score t τ=1 Pτ obtained from Eq. (1). As a participant can select multiple classes at any episode, for the RAM model, instead of predicting a single class based on highest probability, we consider the mean probability over all classes as a threshold and predict the set of classes ct with probabilities greater than the threshold. This ct is used for calculating the score using Eq. (1). Under these conditions, RAM requires 3.7, 8.5, and 7.6 samples to recognize MNIST numerals, uppercase, and lowercase EMNIST alphabets, which correspond to 8.9%, 21.0%, 18.7% of image area respectively. Thus, in comparison to our participants (ref. "Data analysis" section), RAM is less efficient. See Table 2. Results from comparing the fixation maps from RAM and the collected data are shown in Table 3. KL is higher due to its sensitivity to zero values. This implies several locations are sampled by the participants but not by RAM. These experiments can be used as a baseline for evaluating locations sampled by an attention model. 

cistanche-Improve memory2

cistanche benefici - Migliorare Memoria 

Discussioni 

The mcAT paradigm, as used in this paper, has certain points of difference from those that primarily rely on eye movements and gazes to study the mechanisms of object recognition. In the latter, salient parts of the scene attract attention first, followed by saccadic eye movements directing the eye gaze to the salient locations27. Gaze is driven by bottom-up and top-down signals which, together with salience information, form priority maps that guide eye movements for object recognition. Since participants in the present study looked at the static images under free-viewing conditions and with ample time at hand (six minutes for T=12 samplings), they likely engaged in a series of saccadic eye movements or visual reasoning28 to explore the image before clicking on an AOI. These eye movements could have been captured in emAT (using an eye tracker) but not in mcAT. However, these eye movements are affected by mind wandering. While mcAT is also affected by mind wandering29, the effect may be reduced whenever the participants respond after visual reasoning. Since eye movements in response to a stimulus are influenced by the task at hand30, the participants' eye movement patterns were likely influenced by the assigned three-step task at each sampling (ref. "Visual task" section). If an eye tracker had been used, the participants' eye movements to explore the sample would have been intermixed with eye movements to click their chosen classes, which would have complicated the interpretation of the visual exploration of the sample. Clicking the class(es) is a necessary step as it reveals, albeit introspectively, the predicted class(es) in the mind of a participant. It is likely that the gazes immediately before and after the AOI selection-perhaps also aided by fixational eye movements31-contributed the most towards the numeral/alphabet recognition. Indeed, we surmise that participants selected diagnostic areas of the image to distinguish between classes, and those areas likely contain a mixture of bottom-up (e.g., visual contrast) and top-down (numeral/alphabet template) diagnostic information. This is consistent with our finding that participants quickly (within 5 samples on average) distinguished between stimulus classes ostensibly by selecting diagnostic patches.

Table 2. Comparison of efficiency between our participants and the RAM model in terms of the average number of samples required to recognize a numeral/alphabet. The percentage of the image area observed is included in parentheses.

Tabella 2. Confronto di efficienza tra i nostri partecipanti e il modello RAM in termini del numero medio di campioni necessari per riconoscere un alfabeto/alfabetico%numerico. La percentuale della area della immagine osservata è in parentesi.

Table 3. Evaluation of fixation maps from RAM for the stimuli presented in the MTurk experiments averaged over all classes and samplings. Std. dev. are included in parenthesis.


Tabella 3. Valutazione delle mappe di fissazione da RAM per gli stimoli presentati negli esperimenti mediati su tutte le classi e campionamenti. Std. dev. sono inclusi in parentesi.

Conclusioni 

Abbiamo introdotto un set di dati mcAT per riconoscere numeri e alfabeti scritti a mano tramite campionamento sequenziale. Il dato è raccolto da 382 partecipanti presentati con immagini selezionate da dataset benchmark (MNIST, EMNIST). Su media, 169.1 risposte per numero/alfabeto classe sono registrate. Il dato è rigorosamente analizzato per rivelare la efficienza del riconoscimento umano visivo. Il partecipanti osservato solo 12,8% di un immagine per il riconoscimento. Abbiamo proposto un modello di base per prevedere la posizione e classe/e un partecipante sarebbe selezionare al prossimo campionamento. Abbiamo mostrato come le nostre condizioni e per essere utilizzate per valutare un modello di rinforzo basato sull'attenzione in confronto per le prestazioni umane. Questo mcAT dataset, con benefici multipli over eye-tracking dati, riempie un lacuna cruciale in modello basato sull'attenzione ricerca in AI, ML, e altre aree.

Riferimenti 

1. Ranzato, M. A. On learning where to look. arXiv:1405.5488, (2014). 

2. Ba, J., Salakhutdinov, R. R., Grosse, R. B., & Frey, B. J. Apprendimento veglia-sonno ricorrente attenzione modelli. In NIPS, 2593–2601 (2015). 

3. Mnih, V. et al. Modelli ricorrenti di attenzione visiva. In NIPS, 2204–2212 (2014). 

4. Ba, J., Mnih, V., & Kavukcuoglu, K. Multiplo oggetto riconoscimento con attenzione visiva. arXiv:1412.7755 (2014). 

5. Dutta, J. K. & Banerjee, B. Variazione in classificazione accuratezza con numero di scorci. In IJCNN, 447–453 (IEEE, 2017).

6. Larochelle, H. & Hinton, G. E. Apprendimento a combinare foveal scorci con una macchina Boltzmann del terzo ordine. In NIPS, 1243–1251 (2010). 

7. Elsayed, G., Kornblith, S. & Le, Q. V. Saccader: Migliorare la precisione dei modelli di attenzione dura per la visione. In NIPS, 702–714 (2019). 

8. van Birre, R. J. Te fonti di variabilità in saccadico movimenti oculari. J. Neurosci. 27(33), 8757–8770 (2007). 

9. Itti, L. & Baldi, P. Bayesiano sorpresa attrae attenzione umana. Vis. Res. 49(10), 1295–1306 (2009). 

10. Egner, S. et al. Attention and information acquisition: Comparison of mouse-click with eye-movement attention tracking. J. Eye Mov. Res. 11(6), (2018). 

11. Peterson, M. S., Kramer, A. F. & Irwin, D. E. Spostamenti Occulti di Attenzione Precedono Involontari Movimenti Oculari. Percept. Psychophys. 66(3), 398–405 (2004). 

12. Jiang, M. et al. Silicio: Salenza in contesto. In CVPR, 1072–1080 (2015). 

13. Kim, N. W. et al. BubbleView: Un interfaccia per il crowdsourcing immagine importanza mappe e tracciamento attenzione visiva. ACM Trans. Comput. Hum. Interact. 24(5), 1–40 (2017).

14. Sermanet, P., Frome, A. & Real, E. Attention for fine-grained categorization. arXiv:1412.7054 (2014). 

15. Egner, S., Itti, L. & Scheier, C. Confronto modelli attenti con diversi tipi di dati comportamentali. Investig. Ophthalmol. Vis. Sci. 41(4), S39 (2000).

16. Navalpakkam, V. et al. Misurazione e modellazione del comportamento del occhio-mouse in la presenza di pagina non lineare layout. In Proc. Int. Conf. WWW, 953–964 (2013). 

17. Matzen, L. E., Stites, M. C. & Gastelum, Z. N. Studying visual search without an eye tracker: An assessment of artificial foveation. Cogn. Res. Princ. Implic. 6(1), 1–22 (2021). 

18. Tafi, A. P. et al. OCR as a service: An evaluation sperimentale di Google Docs OCR, Tesseract, ABBYY FineReader, and Transym. In Int. Symp. Vis. Comput., 735–746 (Springer, 2016). 

19. Memon, J., Sami, M., Khan, R. A. & Uddin, M. Handwritten optical character recognition (OCR): A comprehensive systematic literature review (SLR). IEEE Access 8, 142642–142668 (2020). 

20. Chaudhuri, A., Mandaviya, K., Badelia, P. & Ghosh, S. K. Sistemi ottici carattere riconoscimento. In Sistemi ottici di riconoscimento caratteri per linguaggi diversi con sof informatica, 9–41 (Springer, 2017). 

21. LeCun, Y. et al. Gradient-based learning applied to document recognition. Proc. IEEE 86(11), 2278–2324 (1998).

22. Cohen, G., Afshar, S., Tapson, J. & van Schaik, A. EMNIST: Un estensione di MNIST a lettere scritte a mano. arXiv:1702.05373, (2017). 

23. Gregor, K., Danihelka, I., Graves, A., Rezende, D. & Wierstra, D. DRAW: A recurrent neural network for image generation. In ICML, 1462–1471 (2015). 

24. Friston, K. Te principio della energia libera: A guida approssimativa al cervello?. Tendenze Cogn. Sci. 13(7), 293–301 (2009). 

25. Mirza, M. B., Adams, R. A., Friston, K. & Parr, T. Introduzione di un modello bayesiano di attenzione selettiva basato su inferenza attiva. Sci. Rep. 9(1), 1–22 (2019). 

26. Bylinskii, Z., Judd, T., Oliva, A., Torralba, A. & Durand, F. Cosa fare diversa valutazione metriche raccontare noi su salienza modelli? IEEE Trans. Modello Anale. Mach. Intell. 41(3), 740–757 (2018). 

27. Itti, L. & Koch, C. Modellazione computazionale della attenzione visiva. Nat. Rev. Neurosci. 2(3), 194–203 (2001).

28. Lamme, V. A. F. Funzioni visive generare consapevole vedere. Front. Psychol., 11, (2020). 

29. da Silva, M. R. D. & Postma, M. Errante menti, erranti topi: Computer mouse tracciamento come un metodo per rilevare mente vagabondare. Comput. Hum. Behav. 112, 106453 (2020). 

30. Schütz, A. C., Braun, D. I. & Gegenfurtner, K. R. Movimenti oculari e percezione: A revisione selettiva. J. Vis. 11(5), 9–9 (2011). 

31. Intoy, J. & Rucci, M. Finely tuned eye movements enhance visual acuity. Nat. Commun. 11(1), 1–11 (2020).

Potrebbe piacerti anche